# Word, PDF and Excel in Prisment: Read Them in Place, Convert Them for Your AI Tools

> Prisment 1.1.0 opens .docx, .pdf, .xlsx and .xlsm in place and converts them to Markdown or CSV in your browser — beside the original, at far fewer AI tokens.

- Canonical: https://prisment.io/blog/word-pdf-excel-to-markdown-and-csv/
- Published: 2026-09-07
- Author: Prisment Team (Documentation & Developer Tooling)
- Publisher: Prisment, https://prisment.io/

**Short answer: Prisment 1.1.0 opens Word, PDF and Excel files in place, and converts them to Markdown or CSV entirely inside your browser tab.** The converted file is created in the same folder as the original — `report.docx` becomes `report.docx.md` — so an AI tool reading that folder picks up the lean text instead of the binary. Nothing is uploaded, there is no account, and it costs nothing.

That last part is the point of the whole release. Teams keep specifications in Word, receive vendor documents as PDF and track numbers in Excel, and then hand the folder to Claude, Cursor or ChatGPT. Those three formats are the ones that arrive at the model either unreadable or expensive — and paying for a document's layout is worse than not sending it at all.

This article covers what each format now does in Prisment, what conversion keeps and what it loses, where the output goes and why, and where the guarantees stop. Prisment is free, needs no sign-up, and the full [privacy policy](https://prisment.io/privacy/) says the same things this article does.

---

## 1. What Prisment opens now, and what it converts to

**Three binary formats gained a viewer and a converter, and three text formats gained support alongside them.** Everything below runs in the browser; nothing has a server step.

| File | Opens as | Converts to | Output name |
| :--- | :--- | :--- | :--- |
| `.docx` | A themed reading view — headings, tables, lists, links, images, footnotes | Markdown | `report.docx.md` |
| `.pdf` | Every page drawn to canvas by pdf.js, with a zoom ladder | Markdown | `spec.pdf.md` |
| `.xlsx`, `.xlsm` | The virtualised data grid, one tab per sheet, with the author's cell fills | CSV | `data.xlsx.csv`, or one file per sheet |
| `.mmd`, `.mermaid` | A live, pan-and-zoom diagram | — | — |
| `.txt` | Plain text, editable and creatable | — | — |
| `.png`, `.jpg`, `.jpeg`, `.gif`, `.webp`, `.svg` | A picture, listed in the file tree like any other document | — | — |

Legacy `.doc` is not supported and is not a recognised extension — only the modern Office XML formats are.

Each binary format has a ceiling on what will be read into memory: **100 MB for a PDF, 50 MB for a Word document or a workbook, 25 MB for an image.** A file over its ceiling is never read at all. It opens as a card naming the file and offering to download it, or to open it in its own application from the desktop app — which is a deliberate choice over a tab that dies quietly.

## 2. Why convert at all: the token argument

**A converted file exists to be read by an AI tool for a fraction of what the original would cost.** That objective, not viewer fidelity, decided every default described below.

A `.docx` and an `.xlsx` are both zip archives full of XML. An agent handed one either cannot read it, or reads it through a tool that unpacks a wall of style definitions, revision identifiers and layout attributes around the sentences you actually wrote. A PDF is worse in a different way: its text layer carries running headers, footers and page numbers on every page, and words hyphenated across line breaks.

So Prisment strips the file down to content and structure and nothing else, and puts the result **in the same folder as the original**. That placement is the feature. An agent pointed at a project folder sees `report.docx` and `report.docx.md` side by side and takes the one it can read. You do not have to change your workflow, tell the agent anything, or move files around.

When a conversion finishes, the notice tells you the new file's size and an approximate token count, so the saving is a number rather than a claim.

## 3. Word to Markdown: what survives, and what does not

**A `.docx` converts to the same GFM dialect Prisment's own editor writes, so the result round-trips in the visual editor rather than being a one-way dump.** Headings, bullet and ordered lists, tables, fenced code from a Code style, `> [!NOTE]` callouts from a Callout or Note style, links, bold, italic, strikethrough, superscript, subscript and footnotes all come across.

What does not survive is stated rather than hidden. Every conversion that loses something starts its Markdown with a `> [!NOTE]` callout listing each category and a count:

| Construct | What happens |
| :--- | :--- |
| Merged table cells | Flattened — the spanning cell's text lands in the top-left, the rest are blank |
| Nested tables | Flattened to their text |
| Equations (OMML), comments, tracked deletions | Dropped and counted |
| Tracked insertions | Kept |
| Text boxes | Their paragraphs stay in the flow at the anchor; only the page position is lost |
| Underline and font colour | Dropped silently — Markdown has no syntax for either |

The viewer is honest about the same boundary. It is a **reading view, not a page-faithful one**: no page layout, no fonts, no colours, no headers or footers. Its chrome says so once, and its "Convert to Markdown" action sits in the same bar.

![Prisment showing Prisment-Intro.docx as a rendered document: the file tree on the left lists it beside a PDF and an Excel workbook, the document heading and body text fill the pane with an embedded screenshot below them, and the top bar carries a Convert to Markdown action, a Download button and a WORD badge.](https://prisment.io/images/blog/binary-formats-word-viewer.png)

## 4. PDF to Markdown: three tiers, and the one it refuses

**How much structure you get back from a PDF depends on how the PDF was made, and Prisment names which of three tiers it used in the loss callout.** This is the honest ceiling of any PDF converter, ours included.

1. **Tagged PDF** — a structure tree is present, so headings, paragraphs, lists and tables are read from it directly. This is what Word, Google Docs and LibreOffice produce by default, and it is the high-fidelity case.
2. **Untagged PDF** — only positioned text runs exist, so structure is reconstructed: lines by vertical clustering, paragraphs by gap and indent, heading levels by font-size rank against the body size, bullets by leading glyph, de-hyphenation at line ends, and running headers, footers and page numbers removed when the same text sits in the same place on most pages. Tables are best-effort from column positions; when column detection was unsure, the callout says so.
3. **Scanned or image-only PDF** — there is no text to extract. Conversion **stops with a message and produces no file**. There is no OCR in this release: Tesseract is over 10 MB of wasm and language data, and it belongs to its own opt-in feature rather than to every page load.

Conversion runs page by page with progress and a cancel button, and stops at 1,000 pages with a message. Cancel leaves nothing behind — no partial document.

The viewer is separate from all of that and has no such ceiling, because pdf.js draws the real page: every page onto a canvas at your display's pixel ratio, a zoom ladder from 50% to 300%, and **each document remembers the zoom you left it at**. It opens at 100%, and the percentage button fits the page to the pane.

## 5. Excel to CSV: values you can trust

**A workbook becomes one RFC 4180 CSV per sheet, carrying the display text you see in Excel rather than the raw storage value.** `15%` reads `15%`, a currency cell keeps its formatting, and a date is written as ISO 8601 — `2026-09-05`, or `2026-09-05T14:30:00` when the cell carries a time — because a locale-ambiguous date in a data file is a bug waiting for someone else's parser.

The rest of the value rules, stated plainly: numbers keep up to 15 significant digits as Excel does; text with a leading zero stays text; booleans are `TRUE` and `FALSE`; errors keep their Excel text such as `#N/A`. **Formulas contribute their cached value only and are never recomputed** — a workbook saved without cached values yields empty cells and a warning naming how many. Empty trailing rows and columns are trimmed, and hidden rows are included.

A single-sheet workbook produces `data.xlsx.csv`; a multi-sheet one produces `data.xlsx.<Sheet>.csv` per sheet, with sheet names reduced to file-safe characters. The parse runs in a worker in dense mode and streams rows to a string builder rather than building an array of arrays, so a very large sheet converts without the interface locking up.

The viewer got a related upgrade in the same release: it now shows the author's **cell fills and font colours**, not just the text. In the sample workbook below, one row is highlighted yellow, and that highlight is the reason the row is interesting — a plain grid showed it as ordinary text.

![Prisment showing transaction-metrics.xlsx in its data grid: nine columns of transaction records with row numbers, one row filled yellow exactly as the workbook author left it, and a top bar carrying Convert to CSV, Search, Download and an EXCEL badge.](https://prisment.io/images/blog/binary-formats-excel-grid.png)

## 6. Where the converted file goes, and how you trigger it

**Conversion is offered in two places, and both create the new document in the original's own folder.** Drop a convertible file and one dialog appears before anything is added, naming the files and offering *Open as is* or *Convert*. Or open the original first and use the **Convert to Markdown** / **Convert to CSV** action in the viewer's chrome bar, which leaves the original open and opens the converted file in a new tab.

Mounting a local folder never raises the dialog. Those files are already on your disk; nothing was handed over, so nothing needs a decision.

Names carry their provenance: `report.docx.md`, not `report.md`. If that name is taken, the next free number is used — `report.docx-2.md` — and **an earlier conversion is never overwritten**. In a mounted local folder the file is written to disk beside its source; in a browser-held workspace it lives under the same path in the app's own storage, and the "All Source Files" ZIP is how it comes out.

## 7. What happens to the pictures

**Before a conversion runs, Prisment inspects the document and — only when it actually contains pictures — asks one question with exactly two answers.** Leaving them out is the lean default; the callout at the top of the Markdown names how many were dropped and what they weighed. Keeping them puts the Markdown and an `images/` folder together in a folder named after the document, linked relatively, which costs an AI tool one short line per picture. Kept pictures become workspace files of their own and show up in the preview.

There is deliberately **no embed option**. Inlining a 200 KB picture as a base64 data URI is roughly 270,000 characters — more tokens than a long specification, for one image. Prisment strips any `data:` URI a converter emits, whatever policy you picked.

"Remember my choice" persists the answer, and Settings → Conversion holds it as Ask each time / Leave them out / Keep in a folder, so a decision made in a hurry is not permanent.

## 8. Where the privacy guarantee applies, and where it stops

**Your documents are never uploaded — that is architectural, not a policy promise.** The Word parser, pdf.js, the spreadsheet reader and the Markdown serializer all execute inside your browser tab; there is no upload endpoint for a document to reach. You can verify it in five minutes with your browser's Network tab, which is the subject of [a separate article](https://prisment.io/blog/client-side-markdown-privacy-explained/).

Two scopes worth being exact about, because a vague claim is worse than none:

- **The guarantee is about your documents, not about site traffic.** The hosted site at prisment.io runs Google Analytics and a conversion tag for our own ad campaigns. No document content, no text and no file names are in any of it.
- **A converted file is a new file on your disk or in your workspace.** If you then hand that folder to a cloud AI tool, the Markdown goes wherever you send it. Prisment's job ends at making the file; where it travels next is your call.

## 9. Converting a file, start to finish

1. **Open Prisment and drop the file.** Drag a `.docx`, `.pdf`, `.xlsx` or `.xlsm` anywhere onto the page, or use the drop card on [/word-to-markdown/](https://prisment.io/word-to-markdown/), [/pdf-to-markdown/](https://prisment.io/pdf-to-markdown/) or [/excel-to-csv/](https://prisment.io/excel-to-csv/) to browse for it.
2. **Choose Convert rather than Open as is.** The dialog names every convertible file in the drop. *Open as is* adds the original only; *Convert* adds the converted document.
3. **Answer the picture question if it appears.** It only appears when the document has images. Tick *Remember my choice* to stop being asked.
4. **Wait, or cancel.** Progress shows pages for a PDF and sheets for a workbook. Cancel leaves no partial file.
5. **Read the result beside the original.** The new file opens in its own tab, in the same folder as its source — which is exactly where an AI tool reading that folder will find it.

On the desktop app the same files also open from Finder or Explorer by double-click, and each viewer can hand the original back to Word, Preview or Excel through **Open in default app**.

---

**The short version:** three formats that used to be invisible in Prisment now open in place, convert locally, and leave a lean text file next to the original for whatever reads the folder next. Start on [the workspace](https://prisment.io/workspace/), or on the converter page for the format you have.

## Questions this article answers

### Can I convert a Word document to Markdown without uploading it anywhere?

Yes. Prisment parses the .docx inside your browser tab and writes the Markdown from there, so the document never crosses the network. There is no upload endpoint, no account and no cost. The converted file is named after the original — report.docx becomes report.docx.md — and is created in the same folder, on disk when that folder is a mounted local folder.

### How accurate is PDF to Markdown conversion?

It depends on how the PDF was made, and Prisment tells you which of three tiers it used. A tagged PDF — what Word, Google Docs and LibreOffice export by default — carries a structure tree, so headings, paragraphs, lists and tables are read from it directly. An untagged PDF has only positioned text, so structure is reconstructed with layout heuristics: heading levels from font-size rank, paragraphs from vertical gaps, de-hyphenation at line ends, and running headers, footers and page numbers removed. A scanned or image-only PDF has no text at all; Prisment stops with a message rather than producing an empty file, because this release contains no OCR.

### What happens to images when I convert a Word or PDF document?

You choose, once, and Prisment can remember the answer. Leave them out is the lean default and the callout at the top of the Markdown names how many were dropped and what they weighed. Keep them puts the Markdown and an images folder together in a folder named after the document, referenced relatively, which costs one short line per picture. Images are never embedded as base64: a single 200 KB picture is roughly 270,000 characters, more tokens than a long specification.

### Why convert a document to Markdown or CSV at all instead of feeding the original to an AI tool?

Because a .docx, .pdf or .xlsx is either unreadable to the tool or expensive to read. A Word file is a zip of XML, a spreadsheet is a zip of XML, and a PDF text layer arrives wrapped in page furniture. A lean Markdown or CSV file carries the same content and structure for a fraction of the tokens, and because Prisment writes it into the same folder as the original, an agent reading that folder picks it up on its own.

### Which file types does Prisment open, and how large can they be?

Markdown and MDX, plain text, HTML, CSV and TSV, JSON, YAML, Mermaid diagram files, Word .docx, PDF, Excel .xlsx and .xlsm, and images. The binary formats have their own size ceilings — 100 MB for a PDF, 50 MB for a Word document or a workbook, 25 MB for an image — and a file over its ceiling is never read into memory; it shows a card offering to download or open the original instead.

### Does Excel to CSV keep formulas?

It keeps their results, not the formulas themselves. Prisment writes each cell as its formatted display text — 15% reads 15%, and a date is written as ISO 8601 — while a formula contributes only the value Excel cached the last time it saved. A workbook saved without cached values produces empty cells and a warning naming how many. Each sheet becomes its own RFC 4180 CSV file.
