PDF Document Structure Visualizer

Visualize a PDF's structure in your browser: headings, paragraphs, lists and images outlined on every page in reading order, plus bookmarks and tags. Export as JSON or Markdown.

PDF

or drag and drop file here

PDF

Fast conversionSecure processingNo registration

What the structure visualizer shows

Upload a PDF and each page is drawn with colored boxes around every block the tool finds: headings at three levels, paragraphs, list items and images. Every box carries a number, and the list beside the page repeats the blocks in that order, so you can see the sequence a converter or a screen reader would follow.

Click a box or a list entry to highlight it in both places, and switch block types on and off to focus on one kind of element. Bookmarks and, for tagged PDFs, the structure tags of the first page are listed as well. If the goal is an editable copy rather than a map, PDF to Word does the conversion.

How blocks are detected

ElementHow it is recognized
HeadingsShort blocks set noticeably larger than the body text. The largest size becomes Heading 1, the next Heading 2, then Heading 3. A single bold line at body size is treated as a minor heading.
ParagraphsLines of the same size that sit close together and overlap horizontally are joined into one block; a larger gap or a change of size starts the next one.
List itemsA line that opens with a bullet, a dash, a number or a letter followed by a period or bracket becomes its own list item.
ImagesEvery raster image drawn on the page, measured from the drawing instructions, so the box matches the placed picture rather than the file inside.
Reading orderBlocks are split by the widest empty gap, vertical gutters before horizontal gaps, so two columns are read column by column beneath a full-width title.

How to visualize a PDF's structure

  1. Select a PDF. The first twenty pages are rendered and analyzed in your browser.
  2. Step through the pages. The legend shows how many blocks of each type were found; click a type to hide or show it.
  3. Click any box to see its text, font size and position in the list on the right.
  4. Download the result as JSON with bounding boxes, or as a Markdown outline that keeps headings, list items and image placeholders.

When the map is useful

Before converting a report, check whether its headings are real headings and whether the columns will be read in the right order; if the map looks wrong, the converted document will too. For scanned documents the map stays empty, because there is no text layer to read, and OCR PDF to Word is the tool to reach for.

The Markdown outline is a quick way to turn a PDF into notes or a table of contents. When the whole text matters rather than its skeleton, PDF to Markdown carries the body over as well.

What to expect

Detection works from geometry and font sizes, not from the author's intent. A large pull quote is labeled a heading, a paragraph with a hanging indent may split in two, and tables appear as several paragraphs or list items rather than as a grid. Vector drawings such as charts are not boxed; only raster images are.

Password-protected PDFs cannot be opened here; use Unlock PDF first. Structure tags are read from the file when present, but the boxes themselves are computed from the page, so an untagged document still gets a full map. To check tagging and archival compliance formally, run the PDF/A validator.

Frequently Asked Questions

Does the analysis happen in my browser?

Yes. Pages are rendered and blocks are detected on your device, and the JSON and Markdown downloads are generated locally. The privacy policy describes how uploads and usage data are handled in general.

Why is a paragraph labeled as a heading?

Headings are recognized by size: a short block set noticeably larger than the body text counts as one. A pull quote, a cover line or a large drop cap meets the same test. The label is a description of how the page is set, not a judgment about what the author meant.

What does the reading order number mean?

It is the sequence in which the blocks would be read if the page were linearized: top to bottom, and column by column where the page has columns. It is the order a converter or a screen reader is likely to follow, which is why a wrong order on the map usually predicts a wrong order in the converted file.

Why does my scanned PDF show no blocks?

A scan is a picture of a page, so the only block the tool can find is the image itself. Run the file through OCR first to add a text layer, then the headings and paragraphs become visible.

Are tables detected?

Not as tables. Cells come through as small paragraphs or list items positioned where the cells are, which still shows the reading order a converter will take. For table extraction use PDF to Excel.

What is in the JSON export?

One entry per analyzed page with its size in points and every block in reading order: type, text, font size, bold flag and a bounding box measured from the top-left corner. The Markdown export keeps only headings, list items, paragraphs and image placeholders.