# Your Best Content Is Trapped in a PDF

**Author:** John Morabito (Founder, /winston)
**Published:** September 20, 2026
**Reading time:** 9 minutes
**Canonical:** https://www.winstondigitalmarketing.com/playbooks/pdfs-and-ai-search-citations/

If the only copy of your spec sheet, your price list, your compliance guide or your menu lives inside a PDF, an assistant asked about it has a much harder time quoting you, and it usually quotes somebody else instead. PDFs are not invisible. Crawlers fetch them, search engines index them, and the text inside most of them is extractable. The problem is that a PDF is a weaker candidate for citation than the same words in an HTML page, for reasons that have nothing to do with whether the file can be read. The fix is usually to give the content a home that can be quoted. You rarely have to delete anything.

## What a PDF actually is to a machine

A PDF is a print format. It describes where marks go on a page. When a crawler extracts text from one, it recovers a stream of characters with reasonable accuracy for a file exported straight out of a word processor, and poor accuracy for anything with two columns, a table, or a header and footer repeated on every page. Those elements interleave into the stream and the result reads like someone shuffled the document.

The bigger problem is structure. In HTML, a heading is a heading because it is marked as one. In a PDF, a heading is a line that happens to be bigger and bold, and to a parser that is a run of larger glyphs. So the single pattern that retrieval systems lean on hardest, a heading that names a question with a paragraph underneath that answers it, does not survive the trip. I wrote about why that pattern matters in [how to write content AI actually cites](https://www.winstondigitalmarketing.com/playbooks/how-to-write-content-ai-cites/), and a PDF strips out the part doing the work.

Then there is the anchor problem. An HTML page gives every section a heading with an id, so a passage can be quoted and linked back to the exact place it came from. A PDF has page numbers that shift whenever the document is regenerated, so the unit of citation is the entire file. If a system is choosing between a passage it can point at precisely and a forty-page document it can only point at generally, it takes the passage.

Last, a PDF is a dead end in your link graph. It sits at the end of a link and rarely links anywhere itself, so it never joins the network of pages that tells a crawler how your site fits together. That mechanism is the subject of [internal linking for AI crawlers](https://www.winstondigitalmarketing.com/playbooks/internal-linking-for-ai-crawlers/), and documents are the part of a site that is most often outside it entirely.

## Find out whether yours are read at all

Do not assume either direction. Check, because the answer is usually more specific than people expect.

- Filter the page report in Search Console to URLs ending in .pdf. You will see which documents are indexed and whether they earn impressions. A document that is indexed with zero impressions across a long window is telling you something.
- Look at your access log for requests to your document paths and see which user agents fetched them. This is the only record of what actually happened rather than what should have happened, and it is the same method as [log file analysis for SEO](https://www.winstondigitalmarketing.com/playbooks/log-file-analysis-for-seo/).
- Copy a distinctive sentence from the middle of one document and search for it in quotation marks. If nothing comes back, the text is probably not extractable, which almost always means the file is a scan.
- Ask an assistant a question one of your documents answers, and look at what it cites. Whatever it quotes instead of you is your real competition for that question.

## Convert it, or keep it as the twin

Every document gets one of two decisions, and which one depends on why it is a PDF in the first place.

Some documents are PDFs by accident. Somebody wrote it in a word processor, exported it, and uploaded the export because that is what the upload field accepted. There is no layout requirement and nothing is lost by publishing the content as a page. Convert those. That is most of them.

Some documents are PDFs for a real reason. A printable spec sheet that goes on a wall, a form somebody fills in by hand, a document whose pagination is cited in a contract, a file a regulator or a distributor expects in that format. Keep those. But keep them as the twin rather than the original: the words live in an HTML page, and the PDF is the convenience copy sitting next to it. The printable version stays printable, and the searchable version stops being a file nobody can quote.

## Give every document a front door

The landing page is where this is won or lost, and most of them are a title and a download button. That gives a retrieval system nothing. A page that earns the citation states what the document is and who it is for, carries the summary in ordinary prose, reproduces the document's own section headings as real headings with a paragraph of substance under each, shows the version and the date, links out to the related pages on your site, and then offers the file.

Write it so the page stands on its own. If somebody reads the landing page and never downloads anything, they should still have got an answer. That is the test, and it is also what makes the page worth citing, because a system quoting you is doing exactly what that reader did.

Menus, price lists and class schedules deserve a specific mention here. A downloadable PDF menu is a fine convenience and a terrible answer. The question "what do they serve" or "what does it cost" gets answered from whatever is in the HTML, and if the HTML says "download our menu" then the answer comes from a delivery platform or a directory that typed your items out. [Schema markup](https://www.winstondigitalmarketing.com/playbooks/schema-markup-for-ai-engines-2026/) on the HTML version compounds that, because it labels the items rather than leaving a parser to infer them.

## Scanned files have nothing to extract

Anything that came from a scanner or a phone camera is a picture of text. Old brochures, signed documents, lab results, catalogs from before the current website. There is no text layer, so extraction returns nothing and the file is effectively blank to everything except a human looking at it.

The test takes five seconds: open the file and try to select a sentence. If you cannot select it, neither can anything else.

Running OCR adds a text layer, and it is worth doing, but treat the output as a draft. OCR on a poor scan produces confident errors, and it produces them in part numbers, dosages, prices and dates, which are the parts somebody will rely on. Proofread it or retype it. Then publish the accurate text as a page and keep the scan as the archival copy.

## Naming, linking and the metadata inside the file

Four small things that cost almost nothing and get skipped constantly.

1. Give the file a URL that reads. A path like /docs/stainless-fastener-spec-2026.pdf describes itself; final_v3_FINAL.pdf does not. For a document with no other descriptive markup, the URL is often the clearest signal attached to it.
2. Make it reachable through real links, not only through a search box or a resources widget that builds its list with JavaScript. A document that can only be found by querying a database is a document nothing will find.
3. Set the title metadata inside the file. A surprising number of documents carry a title like "Microsoft Word - Untitled1" from the export, and that string is what a search engine has to work with when it needs a title.
4. Think hard about gating. If a form in front of the download is a genuine business requirement, fine, but understand that a gated PDF contributes nothing to your visibility. Everything you want to be known for has to be said on the ungated page as well.

## The point

Most companies I look at have their most substantive material sitting in documents. The real detail, the numbers, the process, the parts of the business that would actually persuade somebody, and all of it in a container built for printing. The writing is usually fine. Moving it into pages is one of the cheaper wins available, because nobody has to produce anything new. We do this kind of work as part of our [generative engine optimization service](https://www.winstondigitalmarketing.com/services/generative-engine-optimization/), and it is often the first place we find a client's best material.

## Frequently asked questions

### Can AI search engines read PDFs?

Usually yes, with caveats, and being readable is not the same as being quotable. A PDF exported from a word processor carries a text layer that a crawler can extract, so the words are recoverable. What is not recoverable is the structure. A PDF describes where marks sit on a page rather than what each block of text is, so a line that looks like a heading to you is just a run of larger glyphs to the parser. Multi-column layouts, tables, and a header or footer repeated on every page all interleave into the extracted text and make passages harder to lift out cleanly. A scanned or photographed document is worse again, because the text is stored as an image and there is nothing to extract at all. So the honest answer is that your PDF can probably be read, and the same words in an HTML page are still an easier thing for an assistant to quote and attribute.

### Should I convert my PDFs to HTML pages?

Convert the ones that are PDFs by accident and keep the ones that are PDFs for a reason. If a document exists as a PDF because somebody exported a Word file and uploaded it, that is an accident of workflow and the content belongs in an HTML page. If the fixed layout genuinely matters, because it is a printable spec sheet an engineer pins to a wall, a form somebody fills in by hand, or a document whose pagination is referenced elsewhere, keep the file. The rule I use is that the words live in HTML and the PDF is the convenience copy rather than the original. That way the printable version stays printable and the searchable version is a page with real headings, real internal links, and something an assistant can quote a paragraph out of.

### How do I check whether my PDFs are being crawled?

Three checks, and they take about twenty minutes. First, open Search Console and filter the page report to URLs ending in .pdf, which tells you which documents have been indexed and whether they get impressions. Second, look at your server or CDN access log for requests to your document paths and see which user agents actually fetched them, which is the only place that records what happened rather than what should have happened. Third, copy a distinctive sentence from the middle of one document and search for it in quotation marks; if nothing comes back, the text is very likely not extractable, which usually means the file is a scan. As a fourth informal check, ask an assistant a question your document answers and look at what it cites instead of you.

### What should a PDF landing page contain?

More than a title and a download button, which is what most of them are. A useful landing page states what the document is and who it is for, carries the summary in ordinary prose, and reproduces the document's own headings as real HTML headings with a paragraph of substance under each one. It should show the version and the date, link to related pages on your site, and then offer the file. The point is that the page becomes the thing an assistant can quote and cite, with a stable URL and an anchor for each section, while the PDF stays available for anyone who wants to print or file it. A page that only announces a download gives a retrieval system nothing to work with, so the document stays invisible no matter how good it is.

### Do scanned PDFs work for SEO?

Not on their own. A scanned document is a picture of text, so there is no text layer to extract, and a crawler gets a file it cannot read. The quick test is to open the file in any viewer and try to select a sentence; if you cannot select it, neither can anything else. Running OCR adds a text layer and is worth doing, but OCR quality on a poor scan is unreliable and it will produce confident errors in exactly the places that matter, like part numbers and figures, so anything a customer relies on needs proofreading or retyping rather than trusting the output. Once you have accurate text, the better move is usually to publish it as an HTML page and keep the scan as the archival copy.
