基于 HN 与 Internet Archive 构建 10,000+ 公司招聘板的实时职位搜索
A live job search over 10,000 company career sites, built from public links (HN + the Internet Archive)
作者从 Hacker News「Who is hiring?」评论和 Wayback Machine CDX 索引中提取 Greenhouse、Ashby、Lever、Workday 等 ATS 招聘板链接,验证后得到约 10,800 个活跃公司招聘板、313,000 个在招职位。
Job aggregators are great, but they're copies: scraped, deduplicated, sometimes days old. If you care about a specific slice, like startups and tech companies, you can go to the source instead. Most of these companies post jobs through an applicant tracking system (Greenhouse, Ashby, Lever, Workday…), and those systems have public JSON endpoints that exist so jobs can be shown anywhere.
The hard part isn't fetching jobs. It's knowing which companies exist on which board.
Where do you get a list of company boards?
You can't list "all Greenhouse customers". But companies post their own board links in public places, and one of the best is Hacker News: every month, hundreds of companies reply to "Ask HN: Who is hiring?" with a link to their careers page.
HN has an official search API (by Algolia) that searches comment text:
const r = await (await fetch(
'https://hn.algolia.com/api/v1/search_by_date?query=ashbyhq.com&tags=comment&hitsPerPage=100'
)).json();
Pull the board slugs out with a regex per system:
const PAT = {
greenhouse: /(?:boards|job-boards)(?:\.eu)?\.greenhouse\.io\/([a-z0-9_-]+)/gi,
ashby: /jobs\.ashbyhq\.com\/([a-z0-9_.%-]+)/gi,
lever: /jobs\.(?:eu\.)?lever\.co\/([a-z0-9_.-]+)/gi,
workday: /([\w-]+)\.(wd\d+)\.myworkdayjobs\.com\/(?:[a-z]{2}-[A-Z]{2}\/)?([\w-]+)/gi,
};
The API caps a query at 1,000 hits, so walk back in time with numericFilters=created_at_i<=…, using the timestamp of the last hit as the new upper bound.
This produced about 2,100 candidate boards across 10 systems.
Most of them are dead
Companies get acquired, switch ATS, or stop hiring. So check each one against its public API and keep only boards with at least one open job:
| System | Candidates | Live with jobs |
|---|---|---|
| Greenhouse | 911 | 319 |
| Ashby | 399 | 270 |
| Workday | 234 | 118 |
| Workable | 242 | 47 |
| Lever | 96 | 44 |
| Personio | 53 | 21 |
| Recruitee | 87 | 21 |
| SmartRecruiters | 52 | 16 |
| Rippling | 26 | 13 |
| Teamtailor | 32 | 11 |
880 live boards. A good start, but HN only sees companies that post there.
Scaling up with the Internet Archive
The Wayback Machine has a public CDX index of every URL it has archived, and job boards get archived a lot. You can list archived URLs under a host prefix:
https://web.archive.org/cdx/search/cdx?url=jobs.ashbyhq.com/&matchType=prefix&fl=original&collapse=urlkey&from=2025&limit=50000&showResumeKey=true
Take the first path segment as the slug, page through with resumeKey, and pause a few seconds between requests. It's a shared public service, so be gentle. That gave about 33,000 candidates. After checking each one against its board's API:
| System | Live boards |
|---|---|
| Greenhouse | 5,490 |
| Ashby | 3,642 |
| Lever | 1,453 |
| Workday | 118 |
| Others | ~170 |
About 10,800 live company boards, with 313,000 open jobs. One lesson: Workable started answering 429 halfway through the checks, so I stopped and left it alone. Rate limits are a request, not a challenge.
One snapshot a day, not 10,000 requests per search
With 880 boards you can check every board on every search. With 10,800 you shouldn't: it's slow, and it puts needless load on everyone's job boards. So a small job reads every board once a day and writes one gzipped, line-delimited snapshot (one JSON array per job, about 14 MB for 313k jobs, built in under 5 minutes). Searches stream that file line by line, so a 512 MB container handles it with under 100 MB of memory. My first version parsed the whole file at once, and the container quietly ran out of memory.
Workday is different: big companies have thousands of postings, paged 20 at a time. Downloading 43,000 jobs to find 140 data engineers would be wasteful. Workday's endpoint accepts a searchText, so let it search server-side:
await fetch(`https://${tenant}.${dc}.myworkdayjobs.com/wday/cxs/${tenant}/${site}/jobs`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ appliedFacets: {}, limit: 20, offset: 0, searchText: 'data engineer' }),
});
Then still filter titles locally, because Workday's search also matches descriptions.
A search for "product designer" posted in the last 30 days returns 444 jobs across 10,800 companies in about a minute, most of which is spent on the live Workday searches.
Things that bit me
-
Dates differ by system. Greenhouse gives
first_published, Workday says "Posted 3 Days Ago", and some boards only show the last update. Normalize them, and document which is which. -
"Remote" is inconsistent. Lever and Ashby have structured workplace types; others only have free-text locations. Treat
remoteas "marked remote", not "definitely not remote". - Be polite. One request per board per run, modest concurrency, a User-Agent that says who you are.
Try it
The packaged version is on Apify as Job Search API: title keywords and exclusions, location, remote-only, posted-within-days, and an "only new jobs" mode for daily alerts. You can add your own company boards too. The snapshot refreshes daily, and Workday plus any boards you add are searched live.
What would you add to the list? If your company's board is missing, drop the link in the comments.
Written with AI assistance; the code and numbers were tested before publishing.
来源:Google AI:DEV 作者专属(RSS) · dev.to