← All Tools

exifray

Find the documents a domain publishes and turn their metadata into recon.
GitHub ★ 11
go install github.com/mmarting/exifray@latest
exifray flow diagramSwipe to see the full diagram →

exifray discovers publicly accessible files on a target domain, extracts their metadata in memory and writes nothing to disk. It queries 19 discovery sources, crawls the site and the subdomains that certificate transparency and passive DNS reveal, and refetches from the Wayback Machine anything the live host no longer serves. Where the format allows, only the byte ranges that carry metadata are transferred. It parses EXIF, IPTC, MakerNotes, PDF, OOXML/ODF, OLE2, XMP, ISO base media, PE, RTF and EML, and surfaces findings across 12 categories: usernames, emails, GPS coordinates, internal paths, software versions, printer names, serial numbers, credential-shaped strings, handling markings and superseded revisions. A correlation pass then clusters name spellings into identities, infers the organisation's email format, builds a software inventory with end-of-life notes, profiles authoring hours and time zones, and groups GPS coordinates into sites. Results go to the terminal, JSON, NDJSON, per-category artifact files, Markdown or HTML.

How it works

Discovery
Queries 19 sources concurrently: passive archives, certificate transparency, passive DNS, search engines, sitemaps, robots.txt history and a crawler: to find file URLs on the target. A URL list can be fed in directly with -u to skip this phase entirely
Expansion
Subdomains surfaced by certificate transparency and passive DNS are crawled for documents, so a CT log becomes a document list rather than a list of names
Deduplication
Merges results from every source, normalises scheme, host case and fragments so the same file is not downloaded twice, and filters by extension
Extraction
Fetches each file and parses it in memory, transferring only the byte ranges that carry metadata where the format allows. Files the live host no longer serves are refetched from the Wayback Machine, and nothing is written to disk
Analysis
Scans the extracted metadata across 12 categories, from usernames and GPS to credential-shaped strings and handling markings, scoring each finding for confidence as well as severity
Correlation
Clusters name spellings into identities, infers the organisation's email format and projects it over discovered names, builds a software inventory with end-of-life notes, profiles authoring hours and time zones, and groups GPS coordinates into sites
Output
Presents deduplicated findings grouped by category, with optional JSON, NDJSON streaming, per-category artifact files, a Markdown report and a self-contained HTML report

Features

  • 19 discovery sources (14 keyless, 4 API-based, 1 opt-in)
  • In-memory processing: nothing is ever written to disk
  • Subdomain expansion: hosts from CT logs and passive DNS get crawled for documents
  • Real crawler with configurable depth, JavaScript bundle parsing and common document-directory probing
  • Byte-range fetching: only the ranges carrying metadata are transferred, with full download as fallback
  • Wayback Machine fallback for files the live host no longer serves, preferring the earliest (least sanitised) capture
  • Correlation pass: identity clustering, email-format inference, software inventory with EOL notes, activity profiling and GPS clustering
  • 12 detection categories: GPS, users, emails, software, printers, serials, paths, URLs, hostnames, secrets, classification markings, revisions
  • 13 file type families, including MP4/MOV/HEIC, PE executables, RTF and EML
  • Deep OOXML reading: comment and tracked-change authors, external workbook links, templates, printer DEVMODE, speaker notes, connection strings
  • Deep PDF reading: object-index parser, signature signer, annotation targets, page text and recovered prior revisions
  • IPTC IIM, MakerNotes and PNG eXIf, with stock-agency credits excluded from the people list
  • Confidence scoring on every finding, separate from severity
  • Result cache for extraction and per-source URL lists (172s cold vs 33s warm on a real target)
  • Local directory mode: run every parser over files already on disk, no network at all
  • Diff mode against a previous run, with webhook notification for continuous monitoring
  • Artifact output: users.txt, emails.txt, candidate-emails.txt, software.csv, gps.kml, urls.txt, report.md and report.html
  • Direct URL input from stdin, so gau, katana or waymore can feed it
  • Combined mode merges every domain of an organisation into one people list
  • Proxy support (HTTP, SOCKS5), rate limiting, retries with backoff and JSON/NDJSON output

Discovery Methods

MethodTypeDescription
Wayback MachineIncludedCDX API for historical file URLs, streamed and paged with resumeKey
Common CrawlIncludedCC Index API across several recent indexes, not just the newest
Web ScrapingIncludedCrawls in-scope pages, reads JavaScript bundles and probes common document directories
SitemapIncludedParses sitemap.xml and linked sitemaps, gunzipping compressed ones
robots.txt HistoryIncludedEvery archived version: Disallow paths are a list, written by the target, of what it did not want crawled
CertSpotterIncludedCertificate Transparency logs, keyless tier (a few queries per hour)
crt.shIncludedCertificate Transparency logs; frequently returns 502, CertSpotter covers the same ground
RapidDNSIncludedSubdomain discovery via rapiddns.io, feeding the expansion phase
OTX Passive DNSIncludedHostnames from passive DNS, feeding subdomain expansion
AlienVault OTXIncludedURL list for the domain (optional API key raises rate limits)
URLScan.ioIncludedSearch API paged with search_after (optional API key raises rate limits)
HackerTargetIncludedHost search API for domain-associated URLs
DuckDuckGoIncludedFile-type dorking with no key required
VirusTotalAPI keyURL discovery via the domain endpoint (free tier: 500 lookups/day)
Brave SearchAPI keyFile-type dorking against an independent search index
GitHub Code SearchAPI keyDocument URLs referenced in public repositories
ChaosAPI keyProjectDiscovery subdomain dataset
Public BucketsOpt-inAnonymous listing of S3, GCS and Azure containers named after the target (-s buckets)
Google SearchPaid APIFile-type dorking via Custom Search API (requires API key + Custom Search Engine ID)

Usage Examples

exifray -d example.com
Scan a domain using all keyless sources
exifray -d example.com --json -o results.json
Export findings to a JSON file
exifray -d example.com --output-dir ./loot
Write the artifact files an engagement actually consumes: users.txt, emails.txt, candidate-emails.txt, software.csv, gps.kml, report.md
exifray -d example.com --diff last.json --webhook "$SLACK_URL" --json -o last.json
Continuous monitoring: report only what is new since the last run and push it to Slack
exifray -d example.com --cache-dir ~/.cache/exifray --output-dir ./loot
Cache results so a re-run skips work that already succeeded (172s cold, 33s warm, identical findings)
exifray --dir ./downloaded-docs --output-dir ./loot --html
Extract metadata from files already on disk, with no network at all
exifray -d example.com --min-severity notable --limit 5000
Only the findings that matter, for a large scan
exifray -d example.com --urls-only > docs.txt
Use exifray purely as a document-URL discovery tool
gau example.com | grep -E '\.(pdf|docx|xlsx)$' | exifray -u -
Feed URLs from another recon tool and skip discovery
exifray -l domains.txt --combined --output-dir ./loot
Scan an organisation that owns several domains and get one merged people list
exifray -d example.com --sources wayback,scrape,sitemap
Use only specific discovery sources
exifray -d example.com --archive-only
Work from archived snapshots without touching the live host
exifray -d example.com --proxy http://127.0.0.1:8080 -k
Route through an intercepting proxy such as Burp
subfinder -d example.com -silent | exifray
Pipe domains from subfinder
exifray -q -d target.com | grep "^\[Users\]" | cut -d' ' -f2-
Filter findings by category with standard tools
exifray -d target.com --json | jq '.intel.candidate_emails[]'
Pull the projected email addresses out of the correlation pass with jq

Options

FlagDescription
-d, --domainTarget domain (required unless -l, -u or stdin)
-l, --listFile with a list of domains (batch mode)
-u, --urlsFile with a list of URLs, or - for stdin (skips discovery)
--dirScan files in a local directory instead of the network
-s, --sourcesSources to use, comma-separated (default: all)
-e, --extensionsFile extensions, replacing the defaults
--ext-addFile extensions to add to the defaults
--expand-hostsCrawl this many discovered subdomains, 0 disables (default: 25)
--crawl-depthCrawl depth per host (default: 2)
--crawl-pagesMax pages crawled per host (default: 60)
--crawl-knownVisit N pages other sources found, looking for document links (default: 150)
--cc-indexesCommon Crawl indexes to query (default: 3)
--limitMax file URLs to analyze, 0 = unlimited (default: 0)
--urls-onlyPrint discovered file URLs and exit
-w, --workersConcurrent workers (default: 20)
--host-concurrencyMax simultaneous requests per host (default: 5)
--timeoutHTTP timeout in seconds (default: 15)
--max-retriesMax retries on transient errors (default: 2)
--retry-delayRetry delay in seconds (default: 2)
--max-sizeMax megabytes downloaded per file (default: 50)
--range-fetchFetch only the bytes that carry metadata (default: true)
--archive-fallbackRefetch from the Wayback Machine when a file is gone (default: true)
--no-archive-fallbackDisable the Wayback Machine fallback
--archive-rate-limitMax Wayback requests per second, 0 = unlimited (default: 2)
--archive-onlyFetch from the archive without trying the live host
--archive-latestUse the most recent capture instead of the earliest
--cache-dirCache results here so re-runs skip completed work
--cache-ttlExpire cached entries after N days, 0 = never (default: 0)
--cache-discovery-ttlHours a cached URL list stays usable (default: 24)
--rate-limitMax HTTP requests per second, 0 = unlimited (default: 0)
--proxyProxy URL (http:// or socks5://)
-k, --insecureSkip TLS verification (for an intercepting proxy)
-v, --verboseEnable verbose output
-q, --quietSilent mode: findings only, one per line
--jsonOutput results as JSON
--streamEmit one NDJSON record per result as it completes
-o, --outputWrite results to file (JSON)
--output-dirWrite per-category artifact files and a report
--show-urlsShow source file URLs per finding
--min-severityMinimum severity: info, interesting, notable (default: info)
--diffReport only findings absent from a previous JSON run
--no-intelSkip the correlation pass
--htmlAlso write report.html into --output-dir
--webhookPOST a summary here when findings exist
--combinedMerge results across all scanned domains
--exit-codeExit 1 when nothing is found
-c, --configConfig file path (default: $HOME/.exifray.conf)
--versionPrint version and exit
-h, --helpDisplay help information

This tool covers one part of the job. In a pentest, it sits inside the wider process of mapping the surface, testing trust boundaries and documenting reproducible findings.