Skip to content

About

Conservative HTTP-only technographic detection CLI with evidence-backed Wappalyzer-style fingerprints

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Technograph

Technology intelligence for sales research and GTM workflows.

Technograph scans public websites and DNS records, identifies observable technologies, and returns the evidence behind each finding. Use it to research accounts, inspect integrations, export prospect scans, and track changes through a CLI or local MCP server.

The current candidate supports deterministic quick/deep HTTP collection and an opt-in hybrid Chromium profile. It also turns public DNS assertions into clearly labelled relationships without misreporting them as website technologies.

Quick start

brew install b1rd33/tap/technograph
technograph scan stripe.com shopify.com --format table
technograph explain stripe.com
technograph scan --input domains.txt --format csv --output results.csv

Use a text file with one bare domain per line. JSON is the default output; diagnostics go to stderr.

technograph scan stripe.com --output snapshot.json
printf 'stripe.com\nshopify.com\n' | technograph scan --format jsonl
technograph explain stripe.com --profile deep --pages 3
technograph scan stripe.com --profile hybrid --format table
technograph scan protected.example --profile hybrid --interactive
technograph explore stripe.com shopify.com

--profile quick is the one-page default (8 requests, 2 MiB, 8-second HTTP timeout). --profile deep allows eight pages, four same-host referenced assets, 24 total wire requests, 8 MiB, and a 30-second HTTP timeout. The overall domain deadline is one second beyond the larger HTTP/DNS timeout (normally 9 seconds quick, 31 seconds deep); collection.budgets.deadline_ms records the enforced value. Redirects, robots and retries consume the same aggregate budgets. --pages 1 through --pages 8 can lower the page cap.

--profile hybrid adds one isolated Chromium navigation, rendered DOM/runtime signals, and up to 250 redacted browser resources. --interactive opens the browser visibly and waits for a human to complete a normal challenge; it does not automate CAPTCHA solving. Browser requests pass through Technograph's public-address safety proxy. DNS verification tokens and SPF IPs are redacted unless --include-raw-dns is explicitly supplied.

Installation

Homebrew installs both technograph and technograph-mcp. Upgrade with brew upgrade technograph.

To build from source, use Go 1.25 or newer:

git clone https://github.com/b1rd33/technograph.git
cd technograph
make build

The binaries are written to bin/technograph and bin/technograph-mcp. Prebuilt macOS/Linux ARM64 and AMD64 archives are available from GitHub releases. Release downloads and Homebrew reflect published versions, not unreleased v2 work.

Understand the results

Every input gets an ok, partial, blocked, failed, or invalid status. Findings include stable IDs, matching evidence, roles, confidence class, technology categories, and directly captured versions when available. JSONL streams completed domain results; JSON, table, and CSV preserve input order. JSONL file output streams to a temporary file and replaces the destination only after successful completion; cancellation or output failure preserves the old file.

A detection means that a public signal matched a fingerprint. It does not prove that code executed, a company uses a tool internally, or every part of its stack was discovered. An empty result means no technology was observed under that scan's coverage and catalog.

explain groups findings with their source evidence and collection issues. explore provides a terminal view with list, filter TEXT, show DOMAIN, and clear.

Truncated collection is reported explicitly. Structured reports include http.signals_truncated and http.signals_truncated_reasons; affected otherwise-successful results become partial. Confidence values currently describe individual fingerprint rules, not calibrated probabilities.

See the agent interface and scan schema for exact fields and exit behavior.

Inspect and extend coverage

The v2 embedded catalog contains 70 technologies and 104 patterns. Inspect what the scanner can recognize before interpreting an account's results:

technograph fingerprints
technograph fingerprints --technology hubspot --format table
technograph fingerprints --channel script --channel header
technograph fingerprints --fingerprints custom.json --validate
technograph scan example.com --fingerprints custom.json

An offline importer converts supported native Wappalyzer-style definitions:

technograph fingerprints import --input source/technologies --output generated.json --report compatibility.json

External data needs its own provenance and licensing review. Import success establishes format/regex compatibility, not detection accuracy. Unsupported channels and expressions are reported; no corpus download occurs during a scan.

Read the generated catalog, import guide, and matcher compatibility notes.

Track changes

technograph watch stripe.com --store .technograph-history
technograph history stripe.com --store .technograph-history
technograph compare before.json after.json --output diff.json

watch performs one scan, compares compatible local history, and saves an observation. It is not a background scheduler. --fail-on-change returns exit status 3 after writing results when confirmed changes are reported. Incomplete scans suppress uncertain removals. Observations are time-specific; a missing detection does not establish that a company uninstalled a product.

See automation for scheduling patterns.

Qualify and export accounts

technograph filter --report scan.json --technology Shopify --technology Klaviyo --match all --output qualified.json
technograph filter --report scan.json --technology GA4 --max-age 168h --unknown-output unknown.json --output qualified.json
technograph explain --report qualified.json
technograph export --report qualified.json --format matrix --output qualified.csv
technograph scan --input prospects.txt --profile deep --resume .technograph-resume --output scan.json

Negative filters exclude incomplete accounts as unknown; --without exclusions always apply, including with --match any. Probable matches and stale reports can be exported separately using --unknown-output. Matrix cells are confirmed, probable, not_observed, or unknown; they never turn incomplete coverage into a claim that a technology is absent. Resume identity includes the input, scanner, catalog, policy, profile, page limit and effective deadline. Each completed domain is saved before the rest of the batch finishes. Each result's observed_at records its actual collection start and survives resume. --max-age checks that timestamp, not the report generation date; older reports without it and future-dated observations are treated as unknown.

Connect an agent

{
  "mcpServers": {
    "technograph": { "command": "technograph-mcp" }
  }
}

The stdio server provides scan_domain, scan_domains, explain_domain, validate_domain, and list_fingerprints. Starting it with --history /absolute/path/to/history also enables watch_domain and domain_history. Agents cannot change that storage path.

MCP accepts quick, deep, or non-interactive hybrid profiles, uses embedded fingerprints, and bounds domain batches and concurrent work. CLI and MCP use the same scanner and evidence pipeline.

Collection and access

Technograph uses public HTTP(S), DNS, and opt-in Chromium collection. TLS verification is enabled by default, responses and extracted signals are bounded, and public-address checks apply to connections and redirects. Autonomous requests use ports 80/443 and ignore proxy environment variables.

Cloudflare delivery headers can identify Cloudflare without exposing the origin application. Challenge HTML is excluded from application detection; usable DNS and response metadata remain available. Hybrid interactive mode can pause while the user completes the site's normal challenge. It does not bypass challenges.

For trusted local testing, the CLI has an explicit --allow-private-network option that MCP does not expose. Custom resolvers use repeatable --dns-server IP[:port] arguments. DNS queries can still disclose submitted names to the resolver, including names whose HTTP connections are rejected.

Architecture and development

  • internal/app assembles the shared scanner.
  • internal/domain validates domains and derives registrable apexes.
  • internal/probe collects HTTP and DNS with bounded resource use.
  • internal/extract turns responses into channel-specific signals.
  • internal/fingerprint compiles rules and emits detections with evidence.
  • internal/scanner coordinates concurrent domains.
  • internal/agentapi, internal/agentcli, and internal/mcpserver expose results.
  • internal/history stores observations; internal/policy controls network access.
make fmt
make vet
make test
make race
make benchmark
make build

Tests cover extraction, channel isolation, network failures, resource limits, challenge suppression, deterministic output, history, and MCP integration. See contributing, performance checks, compatibility policy, and v2 benchmark work.

Product direction

The release candidate combines reviewed static evidence with structured DNS relationships and an opt-in browser collector. Later milestones expand reviewed browser fingerprints and relationship-change reporting.

About

Conservative HTTP-only technographic detection CLI with evidence-backed Wappalyzer-style fingerprints

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages