跳到正文
原文
Google AI:DEV 作者专属(RSS)· Clear Fetch·· 9 小时前AI 评分50

Clear Fetch 推出五行 Python 代码调用的 Wappalyzer API 替代工具 tech-stack-detector

A Wappalyzer API alternative in five lines of Python

AI 导读

Clear Fetch 在 Apify 上发布 tech-stack-detector,用五行 Python 代码即可查询网站技术栈,作为 Wappalyzer 和 BuiltWith 按月订阅 API 的按次付费替代。

正文

"What is this website built with?" sits behind a lot of everyday work: qualifying a lead list by the tools companies
already pay for, sizing up competitors, auditing client sites, finding every store on a given platform. Wappalyzer and
BuiltWith answer it, but their APIs come on monthly plans. If you only need a few thousand lookups now and then, paying
per website is simpler.

The five lines

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("clearfetch/tech-stack-detector").call(run_input={"urls": ["stripe.com", "shopify.com", "bbc.co.uk"]})
for site in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(site["url"], site.get("techNames"))

pip install apify-client, and a free Apify account gives you the token. The same call works from Node.js, curl, n8n,
Make or Zapier.

What comes back

One row per website. This is stripe.com from a run today, trimmed:

{
  "url": "https://stripe.com/",
  "ok": true,
  "title": "Stripe | Financial Infrastructure to Grow Your Revenue",
  "techNames": ["Amazon S3", "Cart Functionality", "HSTS", "Nginx", "Open Graph", "Priority Hints"],
  "technologies": [
    { "name": "Amazon S3", "confidence": 100, "categories": ["CDN"] },
    { "name": "Nginx", "confidence": 100, "categories": ["Web servers", "Reverse proxies"] }
  ],
  "dnsTechNames": ["Google Workspace", "Salesforce", "Atlassian Cloud", "Linear", "DocuSign", "HackerOne", "Postman"],
  "scanTimeMs": 1238
}

The dnsTechNames list is the part people underestimate. Companies prove domain ownership to every SaaS tool they set
up by adding a TXT record, and those records are public. Stripe's name its mail provider and a long list of the
software its teams use. For sales prospecting that is often more telling than the website itself.

Every technology also carries its evidence (the header, cookie, script URL or record that matched) and a confidence
score, so you can filter out anything indirect.

Lists of 10,000 domains

  • Paste the list, or put a link to a text or CSV file in urlListUrl. A Google Sheet published as CSV works; header rows and extra columns are skipped.
  • On Apify's default settings it analyzes about six websites a second, so 10,000 take roughly half an hour.
  • If a list is too long for the run's time limit, or reaches the budget you set, the run stops cleanly and saves the websites it did not reach in an UNPROCESSED record, ready to pass as the input of the next run.

How it detects technologies

It fetches each homepage once over plain HTTP, no browser, and reads headers, cookies, HTML and meta tags, script
URLs, inline scripts, the DOM and DNS records. The patterns come from the open
webappanalyzer fingerprints, the continuation of Wappalyzer's public
ones: 7,600+ technologies. It cannot see anything that only appears after JavaScript runs in a browser, unless
another fingerprint implies it.

What it costs

$0.05 per website analyzed ($0.02 until 17 October 2026). Websites that cannot be reached are not charged, there is no
subscription, and paid Apify plans get 10-30% off.

来源:Google AI:DEV 作者专属(RSS) · dev.to