Applying Web Accessibility to a Word Document Editor
What is web accessibility?
Web accessibility means making web information and functionality available to everyone, regardless of disability or device. Screen reader support is often the first thing that comes to mind, but accessibility covers more than that. It also asks whether every feature works with a keyboard, whether the current focus is visible, whether screen readers receive state changes, and whether HTML elements express the right meaning.
Note: A screen reader is software that presents text and an element's role, name, and state through speech or braille. It does not interpret a page by looking at its visual appearance. Instead, it reads the accessibility tree that the browser builds from HTML and ARIA, so identical-looking interfaces can communicate different information depending on their markup.
The Web Content Accessibility Guidelines (WCAG), published by the World Wide Web Consortium (W3C), describe accessibility through four principles:
- Perceivable: People must be able to perceive information and state in different ways
- Operable: People must be able to operate the interface with input methods other than a mouse, including a keyboard
- Understandable: People must be able to understand the interface and its behavior
- Robust: Different browsers and screen readers must be able to interpret the content
What is ARIA?
Accessible Rich Internet Applications (ARIA) is a set of roles, states, and properties that communicates UI information that HTML alone cannot fully express to screen readers.
roleidentifies whether an element is a button, dialog, or status regionaria-labelgives an element an accessible name when visible text is insufficientaria-expandedandaria-disabledcommunicate expanded and disabled statesaria-liveannounces dynamic changes to screen readers
ARIA only communicates an element's role and name to a screen reader. Adding
role="button" to a div does not automatically implement activation with
Enter and Space or disabled behavior. Native HTML should provide the element's
meaning and behavior first, with ARIA filling in information that HTML cannot
express.
Many interfaces use elements such as buttons and form controls, where native HTML already provides roles and keyboard behavior. A Word editor must also communicate document structure and information such as the current paragraph, table dimensions and cell positions, and image alternative text. Mouse-oriented features such as selecting and reordering shapes also need keyboard alternatives.
This post covers the decisions I made while applying web accessibility to a Word editor, focusing on keyboard and screen reader use.
Accessibility in a document editor
On a typical website, people browse information and submit forms. An editor, by contrast, is an authoring tool for creating an output. A Word editor involves constant interaction: moving between paragraphs, selecting text, editing tables and images, and changing the position and stacking order of shapes.
If some editor features are inaccessible, the problem goes beyond missing information. It can prevent someone from authoring the document itself. I therefore needed to examine whether keyboard and screen reader users could perform the same work.
The structure is also more complex than a typical web page. A single screen contains an editing region, toolbar, menus, and dialogs. UI focus and the selected document object are separate states. Because the document is rendered as pages, paragraphs and tables can split into multiple fragments at page boundaries. The HTML element that receives text input is also separate from the HTML elements that display the entered text.
Adding a few ARIA attributes was not enough. I needed to consider the entire flow: finding the editing region, understanding the document structure, confirming the result of navigation, and running mouse-driven features with a keyboard.
The goal was not to complete every aspect of document editor accessibility at once. I first identified the core tasks that needed keyboard and screen reader support, then implemented them incrementally without disrupting the existing editing architecture.
Accessibility requirements
I first examined how the editor represented the document and processed keyboard input. From there, I identified the information and behavior required by keyboard and screen reader users.
Information for screen readers
- The role and name of the document editing region
- The destination and current paragraph after keyboard navigation
- The total row and column count of a table and the current cell position
- Alternative text for images
- Button roles for clickable UI
- Dialog and progress states
- A distinction between meaningful images and decorative icons
Keyboard operations
- Moving, resizing, and reordering shapes
- Keyboard behavior and visible focus for clickable UI
These features depend on each other. Communicating an element's purpose to a screen reader does not make the feature usable if it cannot be operated with a keyboard. Conversely, keyboard navigation is difficult to use when the resulting position is not communicated to a screen reader.
Accessibility information and actual behavior therefore had to be implemented together.
Naming the editing region
The editing region was originally a div used for layout. Its purpose was
visually obvious because it contained the document, but it was difficult to
identify through screen reader landmark navigation.
Note: Landmarks divide a web page into major regions such as the header, navigation, and main content. Screen reader users can open a landmark list and move directly to a region. In an editor, the document can be identified as the main content while toolbars and dialogs receive roles that describe their own purpose. For this work, I made the document editing region a
mainlandmark.
I extended the shared content component to accept optional role and
aria-label props, then provided the appropriate values where the editor uses
it.
<Content role="main" aria-label="Document editing region">
<WordDocument />
</Content>role="main" identifies the page's primary content. The accessible name
“Document editing region” explains its purpose in the landmark list. These
values remain optional because the component is shared by other editors, which
should not all be exposed as a Word document editing region.
The Word editor's input model and existing shortcuts
Because the Word editor renders a paginated document, the HTML element that receives text input is separate from the HTML elements that display document text. Input updates the document's text model, and the resulting state is then rendered onto the pages.
This separation makes the current document position difficult for a screen reader to infer, even when keyboard input works correctly. If paragraph or page navigation only moves the visible caret, the user cannot determine which paragraph or page they reached. The navigation result must be communicated separately.
The editor already implemented these navigation and editing shortcuts. I added
the existing shortcuts to aria-keyshortcuts on the text input element so that
the accessibility tree includes their metadata.
Ctrl + ↑/↓: Move to the previous or next paragraphCtrl + ←/→: Move to the previous or next wordCtrl + Home/End: Move to the beginning or end of the documentCtrl + PageUp/PageDown: Move to the previous or next pageCtrl + F: SearchCtrl + K: Insert a linkCtrl + Alt + M: Add a comment
aria-keyshortcuts does not implement shortcut behavior. It exposes already
implemented key combinations as accessibility metadata. Because screen reader
support and presentation vary, I also provided a separate usage description for
the most important shortcuts. The description includes representative document
and paragraph navigation rather than every shortcut, avoiding a long
announcement before editing begins.
Announcing paragraph navigation to screen readers
Keyboard navigation alone is not enough. After moving, the user needs the new position and content before deciding what to do next.
Note: A live region communicates dynamic page changes to a screen reader without moving focus. When its content changes, the screen reader reads the updated information.
When a shortcut moves between paragraphs or pages, the editor updates the live region with the current paragraph.
Ctrl + ↑/↓: Update the live region with the destination paragraphCtrl + Home/End: Update it with the document boundary and current paragraphCtrl + PageUp/PageDown: Update it with the page navigation result and current paragraph
At the moment of the keyboard event, the document selection can still point to the previous location. I read the current paragraph in the next animation frame, after the edit operation and selection update have completed.
requestAnimationFrame(() => {
const paragraph = getCurrentParagraph();
announce(normalizeParagraph(paragraph));
});I normalized the paragraph text before placing it in the live region. Consecutive whitespace is collapsed, an empty paragraph becomes “Empty paragraph,” and long paragraphs are limited to 200 characters so that navigation is not blocked by a lengthy announcement.
The live region uses role="status" and aria-atomic="true". A screen reader
might not detect an identical string as a change when the user revisits the same
paragraph. To handle this case, the editor clears the live region and restores
the message in the next animation frame.
<p role="status" aria-atomic="true" className="sr-only">
{announcement}
</p>The editor does not update the live region for every keystroke. Doing so would interleave document input with navigation messages and make the content harder to follow. It only sends the current paragraph and navigation details after paragraph or page navigation shortcuts.
Connecting shape reordering to the existing shortcut system
In addition to the navigation and editing shortcuts above, the Word editor already provided keyboard operations for shapes:
- Arrow keys: Move a shape
Shift + Arrow key: Resize a shapeAlt + ←/→: Rotate a shapeTabandShift + Tab: Select the next or previous shapeDelete: Delete a shapeEscape: Clear the shape selection
Instead of creating a separate input system, I added the missing shape reordering behavior to the existing shortcut architecture.
Images, shapes, and text boxes can overlap in a document. A menu or mouse action could move the selected object forward or backward by one level, or send it to the very front or back. There was no equivalent keyboard path.
The chosen shortcuts, Ctrl + ↑/↓ and Ctrl + Shift + ↑/↓, already moved
between paragraphs and extended text selection in the document body. The
operation therefore had to depend on the currently selected element. With only a
shape selected, the shortcut changes the shape's stacking order. With a text
caret or text selection, it retains the existing paragraph navigation or
selection behavior.
const hasOnlyGraphicSelection =
hasGraphicSelection() && !hasBlockSelection() && !hasCaretSelection();
executeCommand(
hasOnlyGraphicSelection ? BRING_FORWARD : MOVE_TO_PREVIOUS_PARAGRAPH,
);The remaining combinations use the same rule to send a shape backward, bring it to the front, or send it to the back. The check covers both shape selection and the presence of a text caret or selection. Without that distinction, shape operations and text editing behavior could conflict.
The keyboard handler calls the existing editing logic instead of changing the
shape's z-index directly. Mouse and keyboard input therefore share the same
domain behavior, avoiding duplicated logic and preserving the existing undo/redo
and document state update flow.
Communicating table and image information
Table position
For a table, its dimensions and the current position matter as much as the cell content. I exposed the total row and column count on the table and the actual position on each row and cell.
- Table:
aria-rowcountandaria-colcount - Row:
aria-rowindex - Cell:
aria-colindex
The cells previously used the gridcell role, but the underlying structure is a
native table rather than an ARIA grid with grid keyboard behavior. A stronger
role is not automatically more accessible. I changed the role to cell so that
it matches the actual structure and interaction.
Image alternative text
The accessible name for an image comes from the non-visual OOXML metadata stored in the document. The editor prefers the author's description, then falls back to the title and name.
I did not expose every graphic as an image. Only Word pictures with a meaningful
description receive role="img" and an accessible name. Pictures without
alternative text remain decorative instead of exposing an internal resource
identifier.
Treating charts, text boxes, and arbitrary shapes as a single image can discard the meaning of their internal text or tables. The implementation therefore considers both the document object type and the presence of a meaningful description rather than assigning the same role to anything that looks like an image.
Improving accessibility in shared UI
An accessible document region is not enough when the toolbar and dialogs remain unusable. I also addressed clear accessibility issues in the shared UI used by Word.
Replacing clickable elements with buttons
Some clickable controls were implemented as div elements with a button role
and keyboard focus. Some also had aria-hidden="true", creating a
contradiction: they could receive focus but did not exist for a screen reader.
I replaced them with <button type="button">. A div is a generic container,
so its button role and keyboard behavior must be implemented manually. A native
button includes Tab focus, Enter and Space activation, and the disabled
state. type="button" also prevents an unintended form submission.
Showing keyboard focus
The global styles removed the default outline, making the focused button
difficult to identify after pressing Tab. I added a 2 px outline to
button:focus-visible so keyboard focus remains visible without showing the
outline for mouse clicks.
button:focus-visible {
outline: 2px solid #005fcc;
outline-offset: 2px;
}Communicating dialogs and status
Dialogs that cover the background and prevent interaction with it receive
aria-modal="true". This tells a screen reader that background content is
unavailable while the dialog is open. Existing focus logic keeps Tab focus
inside the dialog, closes it with Escape, and restores focus to the control that
opened it.
For changing states such as loading, the status region uses role="status" and
aria-live="polite". Updates reach the screen reader without interrupting the
content currently being read. An action such as retrying a load is rendered as a
named button rather than passive status text.
Separating decoration from meaning
Not every visible icon or image needs to reach a screen reader. An arrow next to
a button or another icon that repeats nearby text can add noise. Decorative SVGs
use aria-hidden="true", while decorative PNGs use an empty alt="", removing
both from screen reader navigation.
When an image itself communicates information, such as a warning icon, it can receive a human-readable accessible name. Internal filenames and resource IDs are not exposed as labels.
Menu items now share the same state and behavior across mouse and keyboard
input. Disabled items reject both click and keyboard activation. Enabled items
work with both Enter and Space. Only items with an actual submenu receive
aria-expanded, so the accessibility state matches the available behavior.
What was verified and what remains
The presence of accessibility attributes alone does not prove that a feature is accessible. I separated automated regression coverage, manual checks, and screen reader validation that still remains.
Added automated tests
Shape reordering
The shortcut tests cover both selection branches:
- With only a shape selected, each shortcut runs the corresponding reordering operation.
- Without a shape selection, the existing paragraph navigation and selection behavior remains unchanged.
Shared buttons
The shared button regression tests verify that:
- A clickable control is recognized as a native button.
- The button exposes an accessible name.
- A disabled button does not invoke its click handler.
Manual checks
The following flows still require manual inspection:
- Reach the editing region and shared buttons using only the keyboard
- Inspect the editing region's role and name in the browser accessibility tree
- Confirm live region updates after paragraph and page navigation
- Confirm table dimensions and cell positions
- Distinguish described images from decorative images
- Check focus movement into and out of dialogs
Screen reader validation still required
Automated tests and the browser accessibility tree cannot replace the experience of using a real screen reader. Testing the navigation order and spoken messages with NVDA and VoiceOver across supported browsers remains a follow-up task. Keeping this separate avoids presenting implemented markup as verified usability.
Web accessibility for Canvas
Separate from this Word editor implementation, I also examined how accessibility can be provided when a screen uses Canvas.
The <canvas> itself is a DOM element, but text, shapes, and tables drawn
through the Canvas API do not become individual DOM elements. The browser cannot
infer which pixels represent a paragraph or button or determine their reading
order. A single aria-label on the Canvas therefore cannot communicate all
internal structure and selection state.
Creating a separate HTML accessibility layer
Canvas content needs a separate HTML structure based on the same data so that screen readers can interpret it. The visual output remains on Canvas, while semantic elements such as paragraphs and tables appear in the accessibility tree.
<div>
<canvas aria-label="Document view" />
<section className="sr-only" aria-label="Document content">
<p>Here are the meeting results.</p>
</section>
<div aria-label="Selected shape controls">
<button type="button">Delete selected shape</button>
</div>
</div>Using display: none or aria-hidden="true" removes the HTML from the
accessibility tree. A visually hidden class such as sr-only keeps read-only
content available to screen readers. Interactive elements such as buttons must
remain visible so that keyboard users can see the current focus.
Synchronizing Canvas and accessibility state
Creating the HTML once is not enough. When the selected object or current position changes on Canvas, the corresponding accessibility state must change as well.
- Synchronize the selected paragraph or object with its accessibility element
- Send the current page, selected object, and operation result to a live region
- Provide shortcuts or property controls for features that otherwise require dragging
- Exclude decorative shapes and selection borders from the accessibility structure
The accessibility order should come from the content data model, not screen coordinates. Elements positioned near each other visually do not necessarily follow each other in the logical reading order.
Putting every Canvas object in the Tab order is also inappropriate for a large scene. A list or dedicated navigation model should move between objects and update only the currently selected object's state. Canvas accessibility is therefore not a textual description of pixels. It is a synchronized HTML and keyboard representation of the same underlying data.
Principles from applying web accessibility
I used these principles throughout the accessibility work:
- Prefer native HTML before adding ARIA.
- Provide a keyboard alternative for mouse-driven features.
- Keep visible state and screen reader state consistent.
- Reuse existing editing logic instead of creating separate logic for each input method.
- Do not rely on automated checks alone; verify with a keyboard and real screen readers.
Web accessibility is not a final pass that adds ARIA attributes. It requires reviewing whether people can reach a feature, understand its current state, and produce the same result with different input methods.
A document editor is an authoring tool, not just a place to read information. One inaccessible feature can block part of the authoring process. Visual information and keyboard behavior must also reach screen readers, then be validated through real usage flows.