用 Python 在 5 分钟内通过 API 获取欧洲职位招聘数据
Get European job postings from an API in Python in 5 minutes
开发者用 Job Opportunities API(JOA)写了一个 Python 客户端,可免费获取来自雇主招聘页和 ATS 的欧洲职位数据。该客户端支持按国家、远程状态和发布日期过滤,用游标翻页,并借助字段溯源标签筛选雇主真实公布的薪资,最终导出 CSV。免费 Explore 计划按月记录额度计费,不设试用期。
Disclosure: I build JOA, the jobs API used below. Everything here works on its free plan, and I ran every snippet against the live API in early October 2026 before writing this.
If you have ever needed European job postings as data (for a dashboard, a research project, a job board, or just to see what a market looks like), you know the usual options: scrape career pages yourself, or buy a feed that is a mix of aggregator copies.
This post uses the Job Opportunities API (JOA), which serves postings taken from employers' own career sites and applicant-tracking systems. We will write a small Python client that:
- authenticates with a free key,
- filters by country, remote status and recent date,
- pages through results with the API's cursor,
- uses the per-field provenance tags to ask for salaries the employer actually published,
- writes a CSV, and
- stays inside the free record allowance.
Step 0 - look before you sign up
The statistics endpoints need no key and no account. This gives you the shape of the data (coverage per country, freshness, which sources feed it):
curl https://api.jobopportunitiesapi.org/public/coverage
curl https://api.jobopportunitiesapi.org/public/coverage/countries
Coverage is not uniform across countries, and the report says so openly, so check the countries you care about before building on them. The numbers are live, which is why I am not quoting any here: the always-current version is on the facts page.
Step 1 - get a free key
Sign in with your email at jobopportunitiesapi.org/login, click the link it sends, and press Create a free key on the dashboard. There is no password and no card. The free Explore plan has a monthly record allowance rather than a trial period, so it does not expire (plans and limits).
export JOA_API_KEY='paste-your-key-here'
curl -s -H "Authorization: Bearer $JOA_API_KEY" https://api.jobopportunitiesapi.org/v1/me
/v1/me tells you the plan, the rate limits and today's usage, and it costs no records. Useful to know early: one row returned = one record, so while exploring, keep limit small.
Step 2 - the client
Save this as joa_europe_jobs.py (it needs only requests):
"""Pull European job postings from the Job Opportunities API (JOA) with Python.
Needs: Python 3.9+, `pip install requests`, and a free Explore key in JOA_API_KEY.
"""
import csv
import os
import random
import sys
import time
from datetime import date, timedelta
import requests
API = "https://api.jobopportunitiesapi.org"
# EU27 + EFTA + UK as ISO-3166 alpha-2 codes; the API accepts a comma-separated list.
EUROPE = (
"AT,BE,BG,HR,CY,CZ,DK,EE,FI,FR,DE,GR,HU,IE,IT,LV,LT,LU,MT,NL,PL,PT,RO,SK,SI,ES,SE,"
"IS,LI,NO,CH,GB"
)
session = requests.Session()
session.headers["Authorization"] = f"Bearer {os.environ['JOA_API_KEY']}"
session.headers["Accept"] = "application/json"
def get(path, **params):
"""GET with the retry rule from the docs: retry 429/503 after Retry-After, never retry 402."""
for attempt in range(4):
resp = session.get(API + path, params=params, timeout=30)
if resp.status_code in (429, 503):
wait = int(resp.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait + random.random()) # jitter
continue
if resp.status_code == 402:
sys.exit("This month's free records are spent. Wait for the 1st or upgrade.")
if not resp.ok: # the API explains itself: {"error": ..., "message": ...}
sys.exit(f"{resp.status_code} {resp.json().get('message', resp.text)}")
return resp
sys.exit("Still rate limited after 4 tries.")
def iter_jobs(max_rows=25, page_size=25, **filters):
"""Walk /v1/jobs with the keyset cursor, stopping after max_rows (each row is one record)."""
cursor, seen = None, 0
while seen < max_rows:
params = dict(filters, limit=min(page_size, max_rows - seen))
if cursor:
params["cursor"] = cursor
resp = get("/v1/jobs", **params)
body = resp.json()
for job in body["data"]:
yield job
seen += 1
if not body["has_more"]:
break
cursor = body["next_cursor"]
left = resp.headers.get("x-ratelimit-records-remaining")
print(f" (records left this month: {left})", file=sys.stderr)
def provenance(job, field):
"""Where did this field come from? published / inferred / absent."""
return job["field_sources"].get(field, "absent")
if __name__ == "__main__":
me = get("/v1/me").json() # costs no records
print(f"plan={me['plan']} per_minute={me['limits']['per_minute']} per_day={me['limits']['per_day']}")
week_ago = (date.today() - timedelta(days=7)).isoformat()
print("\n== Remote data-engineering roles in Europe, last 7 days ==")
rows = []
for job in iter_jobs(
max_rows=10,
country=EUROPE,
remote="remote",
title="data engineer",
posted_after=week_ago,
):
rows.append(job)
print(f"{job['country']} {job['title'][:48]:<48} {job['company'][:28]:<28} "
f"remote={provenance(job, 'remote')}")
print("\n== Roles where the employer itself published the salary ==")
for job in iter_jobs(
max_rows=6,
country="DE,FR,NL,IE,ES,IT",
has_salary="true",
require_fields="salary",
title="engineer",
):
lo, hi = job["salary_min"], job["salary_max"]
print(f"{job['country']} {job['title'][:44]:<44} {lo}-{hi} {job['salary_currency']}/{job['salary_period']}")
with open("european_jobs.csv", "w", newline="", encoding="utf-8") as fh:
writer = csv.writer(fh)
writer.writerow(["title", "company", "country", "city", "remote", "posted_at", "apply_url"])
for job in rows:
writer.writerow([job["title"], job["company"], job["country"], job.get("city", ""),
job["remote"], job["posted_at"], job["apply_url"]])
print(f"\nwrote european_jobs.csv with {len(rows)} rows")
A few choices worth explaining:
-
Cursor pagination.
/v1/jobsis keyset-paginated: every response carriesnext_cursorandhas_more, and you pass the cursor back. There are no offsets, so rows that appear while you page can't shift your window. -
Retries follow the docs. Retry 429 and 503 after
Retry-After(with jitter), and never retry 402, which means the month's records are spent. -
Errors explain themselves. The API replies with
{"error": ..., "message": ...}, and the client prints the message. I found this out the hard way, see the gotcha below.
Step 3 - run it
pip install requests
python joa_europe_jobs.py
This is the real output from my run (yours will differ, since the data moves every day):
plan=explore per_minute=30 per_day=5000
== Remote data-engineering roles in Europe, last 7 days ==
FR Data Engineer - Transport - Nantes Sopra Steria remote=published
EE Senior Backend Engineer - Data Core (remote, Eur Modash remote=published
FR Senior Software Engineer - Data Search (remote, Modash remote=published
HR Data / DevOps Engineer (m/f) solvership remote=inferred
DE Senior Data Platform Engineer (m/w/d) | Snowflak Jobrad Loop remote=inferred
FR Finance Data Engineer Voodoo remote=published
FR Stage - Data Engineer - Aeroline - Toulouse Sopra Steria remote=inferred
== Roles where the employer itself published the salary ==
IT Field Automation Engineer 28000-35000 EUR/year
IT Junior Digital Verification Engineer 35000-40000 EUR/year
DE QA Engineer (gn) 55000-75000 EUR/year
IT SALES ENGINEER (m/f/d) 43000-50000 EUR/year
IT PROGRAMMATORE PLC / AUTOMATION SOFTWARE ENGI 35000-45000 EUR/year
IT Data Engineer – Cloud & AI Solutions 33000-40000 EUR/year
wrote european_jobs.csv with 7 rows
Provenance: why remote=published matters
Every row has a field_sources object that tags each field as published (the employer stated it), inferred (derived by the ledger, for example from the title or location) or absent. Look at the first query above: some roles are remote=published, others remote=inferred. If you are building something where a wrong "remote" label is costly, you can now filter on it, and the second query shows the server-side version of the same idea: require_fields=salary returns only rows where the salary came from the employer, not from a guess.
Gotchas I hit while testing
-
require_fieldsaccepts a fixed set of fields (description,employment_type,location,posted_at,remote,salary,source_type). My first attempt usedsalary_currencyand got a clear422 bad_require_fieldslisting the valid ones. That is why the client prints the API's message. -
title=searches the job title only.q=searches title, company and location, anddescription_contains=searches the advert body. Pick the narrowest one. - Structured salaries are a small share of postings in most markets, so a salary-filtered query returns far fewer rows than an unfiltered one. That is expected, not a bug.
- This is not a census of every vacancy. Employers with public career pages are better represented than those without.
Where to go next
- Roles that have left their source are kept with a closure date and reason:
/v1/jobs/closed. I use it in a follow-up post on tracking which companies are hiring and which roles are closing. - If you work with AI agents, there is an MCP connector, so an agent can query the same data without you writing a client.
- The full parameter list is generated from the OpenAPI spec: docs.
- When 1,000 records a month is not enough, paid plans start at EUR 29/month: pricing.
If you hit a rough edge, tell me in the comments. It is one person's product and I read everything.
来源:Google AI:DEV 作者专属(RSS) · dev.to