← all articles

What a document carries when you send it

The file is not just the file

Open a Word document and you see a page of text. Open the same file with a metadata viewer and you see something else: a username, a company name pulled from your software license, a list of who edited the file and when, sometimes a full revision history of text that was deleted and never meant to leave your machine. None of this is a bug. It’s how office software has worked for decades, because metadata is genuinely useful when you’re not thinking about who else will see the file.

This matters more than it used to, because documents move further than they used to. A resume goes to a recruiter, then to an applicant tracking system, then possibly to a public job board. A PDF report gets forwarded three times before someone posts it online. A photo attached to a support ticket ends up in a public bug tracker. At each hop, the file usually still has everything it had when you created it.

The goal here isn’t to make you paranoid about every file you touch. Most metadata is harmless. The point is to understand what’s actually in a document before you send it somewhere the metadata could matter, so you can make that call deliberately instead of by accident.

Where the data actually comes from

Metadata isn’t hidden by some conspiracy. It’s a normal part of how these file formats are built, and each piece has a mundane origin.

Office documents (.docx, .xlsx, .pptx) are ZIP archives containing XML files. One of those XML files, core.xml, stores the author name your copy of Office was registered under, the company name from your license, creation and last-modified timestamps, and how many times the document has been revised. If your organization uses track changes or comments, that content is stored in the document too, even after you accept all changes, unless you specifically clear it. Some versions of Office also keep a “fast save” cache that can include fragments of previously deleted text.

PDFs carry a similar metadata block, usually written by whatever tool created the PDF: Word’s PDF exporter, Adobe Acrobat, a scanner’s software, or a print driver. This includes author, producer, and creation software, plus optional custom fields. PDFs can also retain hidden layers, embedded fonts, form field data, and in some cases the original source file was scanned or exported from, which is sometimes visible in the “Producer” or “Creator” fields.

Images carry EXIF data, written by the camera or phone that took the photo. This can include the camera model, exposure settings, the date and time, and if location services were enabled, GPS coordinates accurate to a few meters. Screenshots typically don’t have GPS data because there’s no camera sensor involved, but they can still carry device and software identifiers depending on the OS.

Spreadsheets often carry more than people expect, because deleted rows, hidden sheets, and old cached values from formulas can persist in the file even after the visible data looks clean. A hidden sheet is still part of the file; hiding it in the UI doesn’t remove it from the underlying XML.

None of this is written to spy on you. It’s written because it’s useful for version control, collaboration, and organizing files, right up until the file leaves the group of people it was useful for.

Why this is a real threat model, not a hypothetical one

The realistic risk isn’t a shadowy adversary decoding your files. It’s much more mundane: metadata revealing something you didn’t mean to share, to an audience you didn’t choose.

A common example is authorship metadata contradicting a public claim. If a document is presented as coming from one person or organization, but the metadata field lists a different author or company name from whoever’s software license created it, that’s an easy inconsistency for anyone who bothers to check.

Another common one is location data on images. A photo posted publicly with GPS EXIF data intact tells anyone who downloads the original file roughly where it was taken. Most social platforms strip this automatically on upload, but direct email attachments, cloud drive links, and messaging apps often don’t.

Revision history is the one people are most often caught off guard by. Comments and tracked changes left in a “clean-looking” Word document are still there in the file structure even if the document looks finished on screen. If that document gets forwarded to someone outside the original group, they can see internal notes, earlier drafts of a negotiating position, or names of people who weren’t supposed to be visible.

None of these examples require sophisticated tools to exploit. Basic metadata viewers are free, built into some operating systems, or a right-click away in most file managers. Treating metadata as low-risk because “nobody’s going to dig for it” undersells how easy digging actually is.

Checking what’s in a file before you send it

The most direct way to know what a document carries is to look. This doesn’t require specialized forensic software.

On Windows, right-clicking a file and opening Properties, then the Details tab, shows a chunk of the metadata for Office documents and images directly, including author, company, and for photos, a GPS section if location data is present. On Mac, Finder’s Get Info panel shows some of this, and Preview’s Inspector shows EXIF data for photos.

Microsoft Office has a built-in tool for exactly this purpose: Document Inspector, found under File > Info > Check for Issues. It scans for comments, tracked changes, hidden text, document properties, and custom XML data, and lets you remove categories selectively. It’s worth running before sending any document that’s been through multiple rounds of edits or has ever had track changes turned on.

For PDFs, most PDF editors have a document properties panel that shows and lets you edit the metadata fields directly. Some PDFs also carry an XMP metadata block, which is a separate, more detailed metadata format layered on top of the basic fields, and not every PDF tool shows both.

For a broader look across file types, ExifTool is a widely used command line utility that reads and can strip metadata from images, PDFs, and many document formats in one pass. It’s not a privacy product, it’s a metadata utility, but it’s precise about what’s actually in a file, which is the useful part.

What removing metadata does and doesn’t change

Stripping metadata removes what’s embedded in the file itself: author names, GPS coordinates, revision history, hidden text, thumbnails. It’s a meaningful step for controlling what a specific file reveals about its own history.

It doesn’t change anything about how the file was transmitted. Email headers, cloud storage logs, and the metadata your email provider or file host keeps about who uploaded or sent what are separate systems, untouched by cleaning the document itself. A stripped document sent from an account tied to your name is still tied to your name through that account, just not through the file’s internal properties.

It also doesn’t retroactively fix copies that already went out. If a document with tracked changes was already sent to five people before you noticed, stripping metadata from your master copy doesn’t touch those five copies already in other inboxes.

And no single check or tool guarantees a document is now “clean.” Different formats hide data in different places, some metadata is regenerated by the software that opens a file next, and it’s easy to check one field and miss another. The realistic goal is reducing what you’re inadvertently disclosing, not achieving some final, verified state of having none.

Building the habit

The practical version of all this is a five-minute check before sending anything that’s been edited more than once, contains images, or is going somewhere outside a small trusted group: open the document properties, run Document Inspector if it’s Office, glance at EXIF if it’s a photo. It’s a habit, not a one-time fix, because every new draft and every new photo generates fresh metadata.

Understanding what’s in your files before you hit send puts the decision in your hands, which is the whole point.

For more breakdowns of how the everyday tools you use actually work, and what that means for what you’re sharing without meaning to, head back to our home page.

from the team
Want a real mobile IP, not a datacenter VPN endpoint?

Shared VPN exit nodes get flagged and blocked. Singapore Mobile Proxy runs real 4G/5G mobile IPs that give you a residential-grade address carriers still trust.

see how it works →
read on
More from The Privacy Wire

VPN and tool reviews, realistic opsec guides, and privacy news for people who want to protect their data.

browse all articles →