The best recon shows up where nobody is looking.
A few years ago that was JavaScript. I wrote about not skipping the bundles back when most people still did, and about what to grep for in the first ten minutes. That edge is mostly gone now. Every methodology post mentions JS, every toolchain pulls it, and the obvious findings in there have been picked over.
The place almost nobody looks today is the documents.
Every organization publishes files. Annual reports, price lists, decks exported to PDF, technical sheets, the spreadsheet somebody linked from a landing page four years ago and never took down. They are public, they are in scope, and hardly anyone opens them.
Not because they are hard to find. Because there is no vulnerability inside them, so nothing in the toolchain reports one. Burp does not flag a docx. There is no template for a PowerPoint. The file returns 200, it is a document, the scanner moves on. A human working from a checklist does the same, because "read the PDFs" is not on the checklist.
Here is what is actually in there:
- Author names, and the account name of whoever last saved the file
- Internal paths, including UNC paths pointing straight at a file server
- Printer models, and the queue a document was sent to
- The exact version of the software that produced it, which is a software inventory nobody had to ask for
- GPS coordinates on photos
- Email addresses in fields nobody thinks of as fields
- In spreadsheets, the connection string the workbook pulls from: server, instance, catalogue, and often the account it connects as
None of that is a vulnerability. All of it is the shape of the organization. Who works there, what they run, how the network is laid out, what the username format is. That is what you want in hand before you touch the application, and it is sitting in files the target published on purpose.
It survives because scrubbing is a process, and processes have gaps. Somebody exports a PDF from Word at eleven at night before a deadline. A contractor sends over a deck that never went through the pipeline. The policy covers the press releases and nobody wrote it down for the technical sheets.
exifray is the tool I use for this, and the reason it exists is that the interesting half of the problem is not the part everybody assumes.
Extraction is the easy half. exiftool has been reading metadata out of files for twenty years and handles more formats than I ever will. The hard half is that you need the files first, and a target's documents are not in one place. They are on subdomains you have not enumerated, in archive snapshots of pages that no longer exist, behind robots.txt entries somebody removed in 2019, in buckets named after the company, linked from repositories nobody remembers publishing.
So most of exifray is not a parser. It is nineteen discovery sources feeding one.
Two things surprised me once I started measuring instead of assuming.
The first is that files which 404 today are not gone. Passive sources return URLs that were live years ago and most of them are dead on the live host. Counting those as unreachable throws away the best part of the corpus, because a document is least sanitized the first time an archive saw it. Refetching from the Wayback Machine, preferring the earliest capture rather than the most recent, changes what a scan returns.
The second is worse, and it was my own fault. Two of the twelve sources I originally shipped were contributing nothing at all. Certificate transparency and passive DNS return host roots, and the extension filter dropped every single one before it reached the fetcher. I did not notice until I measured each source on its own against the same two domains. The fix was not a better parser. It was to treat those hosts as a list to go and crawl rather than a list of names to print.
That is the general shape of this work. The gains are not in the clever part.
A list of names is not intelligence either. Forty documents hand you the same person spelled four different ways: J. Smith, jsmith, John Smith, Smith, John. Cluster those into identities, count what each one authored and date it, and you get something else entirely. Who is still there. Who stopped appearing in the corpus years ago. Which department a shared template belongs to, because a template lives on a share owned by a team.
Then infer the email format from the addresses you did find and project it over the names you did not. The people list becomes a contact list. That is the step where metadata turns into recon.
It does not always work, and the way it fails is informative on its own. A site that resizes images on demand re-encodes them, and re-encoding strips everything. On one target, 31 of 32 files came through the framework's image pipeline and every one of them was clean. That looks like a failed scan. It is not. It is a fact about the target: better discovery will not help, and the effort belongs on documents rather than images.
The same goes for a target that strips consistently. Good metadata hygiene across a whole corpus tells you something about their maturity, and the handful of files that still carry a name are the ones that escaped the process. That makes them more interesting, not less.
One more thing worth measuring: the fields are not the only place. Running the same detectors over the text of the pages, on a sixty file corpus, took findings from 50 to 64. It recovered seven private IP addresses and six email addresses that no metadata field contained. Somebody had pasted them into the body of the document.
If you are on the other side of this, the fix is not a tool. It is a step in whatever process publishes files to your site, and a decision about what public actually means for a document. The cheap version is to run this against your own domain before somebody else does. Twenty minutes, and the output is a list of your own staff, your internal hostnames and your software inventory, assembled entirely from files you published on purpose.
Version 1.2.0 is out today. The changelog is long, and most of it is things that turned out to be wrong once I stopped assuming and started measuring them.
The argument has not changed. The best recon is where nobody is looking. Right now that is the documents.
