What the structure visualizer shows
Upload a PDF and each page is drawn with colored boxes around every block the tool finds: headings at three levels, paragraphs, list items and images. Every box carries a number, and the list beside the page repeats the blocks in that order, so you can see the sequence a converter or a screen reader would follow.
Click a box or a list entry to highlight it in both places, and switch block types on and off to focus on one kind of element. Bookmarks and, for tagged PDFs, the structure tags of the first page are listed as well. If the goal is an editable copy rather than a map, PDF to Word does the conversion.
How blocks are detected
| Element | How it is recognized |
|---|---|
| Headings | Short blocks set noticeably larger than the body text. The largest size becomes Heading 1, the next Heading 2, then Heading 3. A single bold line at body size is treated as a minor heading. |
| Paragraphs | Lines of the same size that sit close together and overlap horizontally are joined into one block; a larger gap or a change of size starts the next one. |
| List items | A line that opens with a bullet, a dash, a number or a letter followed by a period or bracket becomes its own list item. |
| Images | Every raster image drawn on the page, measured from the drawing instructions, so the box matches the placed picture rather than the file inside. |
| Reading order | Blocks are split by the widest empty gap, vertical gutters before horizontal gaps, so two columns are read column by column beneath a full-width title. |
How to visualize a PDF's structure
- Select a PDF. The first twenty pages are rendered and analyzed in your browser.
- Step through the pages. The legend shows how many blocks of each type were found; click a type to hide or show it.
- Click any box to see its text, font size and position in the list on the right.
- Download the result as JSON with bounding boxes, or as a Markdown outline that keeps headings, list items and image placeholders.
When the map is useful
Before converting a report, check whether its headings are real headings and whether the columns will be read in the right order; if the map looks wrong, the converted document will too. For scanned documents the map stays empty, because there is no text layer to read, and OCR PDF to Word is the tool to reach for.
The Markdown outline is a quick way to turn a PDF into notes or a table of contents. When the whole text matters rather than its skeleton, PDF to Markdown carries the body over as well.
What to expect
Detection works from geometry and font sizes, not from the author's intent. A large pull quote is labeled a heading, a paragraph with a hanging indent may split in two, and tables appear as several paragraphs or list items rather than as a grid. Vector drawings such as charts are not boxed; only raster images are.
Password-protected PDFs cannot be opened here; use Unlock PDF first. Structure tags are read from the file when present, but the boxes themselves are computed from the page, so an untagged document still gets a full map. To check tagging and archival compliance formally, run the PDF/A validator.