跳到正文
原文
Google AI:DEV 作者专属(RSS)· Artem Lazarev·· 10 小时前AI 评分54

用 Apify MCP connector 让 YC Jobs Scraper 免持 Notion token 直写 Notion

I gave my YC jobs scraper write access to Notion without ever holding a Notion token

AI 导读

作者为 Apify Store 上的 YC Jobs Scraper 接入 Apify MCP connector,通过 Apify 托管代理以用户授权的 Notion 账号写入数据,Actor 全程接触不到 Notion 凭据,省去 CSV 导出步骤。

正文

TL;DR

Apify MCP connectors use the Model Context Protocol (MCP) to let an Actor write into a user's connected apps without ever touching that app's credentials. I used one to push my scraped Y Combinator jobs straight into Notion and delete my CSV export step. Below: why a token field was never an option for a published Actor, the documented Python snippet that couldn't run on any SDK version, and five steps covering connector setup, tool discovery, field mapping, safe writes, and what the whole thing costs per run.

The workflow I wanted

My Y Combinator jobs scraper has run happily for months. It walks the YC directory, pulls every company that's hiring, and returns salaries, equity ranges, founder profiles, and full job descriptions. The last full sweep did 1,429 companies and 3,305 jobs in about an hour and fifty minutes.

Then the data sat there. An 18 MB JSON file in a dataset. What I actually wanted was boring and specific: a running list of who's hiring, at what salary, in which batch, that I could sort and annotate and come back to a week later. So every time, I exported a CSV, opened it in a spreadsheet, and filtered by hand. The scraper was finished. My workflow wasn't.

Apify MCP connectors closed that gap. They let an Actor call an external service's allowed tools through an Apify-managed proxy, using the user's own authorized account.

The Actor is published on Apify Store, charges per result, and has 120 users I've never met.

After I added connector support, it writes scraped jobs into a Notion database on its own. All it took was a connector field in the input schema, a declaration of the Notion tools it needs, and the code that turns scraped jobs into Notion pages.

The full implementation is in the repo: github.com/artem-lazarev/apify-ycombinator-jobs-scraper.

For users of Apify Actors. To use an MCP connector, the Actor's developer has to add support for it first. Once that's in place, open Settings > API & Integrations in your own Apify account and authorize the connector you need, such as Notion. Then open the Actor's input form, select that connector, and fill in the destination details. For my YC Jobs Scraper, that's the Notion data source ID. Run it, and the scraped jobs land in your Notion database.

Why I couldn't just ask for a Notion token

The obvious option was to add a Notion API token field to the input schema and call the Notion API directly. A Notion integration token is a single secret that grants access to everything you've shared with that integration: pages, databases, all of it.

Asking those users to paste a Notion token into my Actor means asking them to trust that my code won't log it or send it elsewhere. I wouldn't paste my own token into someone else's Actor, so I don't think it's reasonable to ask.

MCP connectors invert that. The user authorizes Notion once, in their Apify account. My Actor receives a connector ID, which is just an opaque string. At runtime it authenticates to the Apify MCP proxy with the run's own Apify token, and the proxy adds the Notion credential server-side before forwarding anything upstream. That credential never enters my container. There's nothing for me to leak and nothing anyone has to trust me about.

Users still need to check which tools an Actor declares, because it can read and write through them. Apify treats the Actor runtime as untrusted precisely because it runs developer-supplied code. I'm the untrusted party here, and I'd rather be.

What you need to follow along

By the end of this walkthrough your Actor will finish a scrape and push the results into Notion on its own. No CSV export, no import step, no Notion token anywhere in your code.

Before any of that works, one thing has to be true: your Notion database has to exist already, with the columns you actually want. This Actor creates pages. It doesn't create databases and it doesn't create columns. Step 3 covers what happens when you gloss over that, which in my case was a batch of jobs landing in Notion with nothing in them but titles.

So build the destination first. Start with a title column and add the fields you plan to fill.

The rest of the list is short:

  • An Apify account.
  • A Python Actor you can build and run in Apify Console.
  • A Notion workspace you can authorize. Apify provides the managed OAuth client for Notion, so you don't need to register your own OAuth app.
  • apify-cli 1.8.0 or newer, if you deploy from the command line.

The connection code targets Python 3.10 or newer, the Apify SDK, and mcp 2.x. One packaging detail worth knowing up front: mcp 2.x depends on httpx2, which is a separate package from httpx. Install and import that one.

The changes land in three places: inputs go in .actor/input_schema.json, the Notion helpers live in src/notion_sync.py, and the sync gets called after results are saved in src/main.py. The repository has the complete scraper; the excerpts below are the integration only.

1. Declare the Apify MCP connector and its allowed tools

When I opened my authorized Notion connector, it offered 28 tools: notion-search, notion-update-page, notion-create-database, and 25 more. My Actor can call two of them. The proxy rejects everything else before it ever reaches Notion.

Setting up that allowlist is what this step is for, and it lives in .actor/input_schema.json. Here's the notionConnector field inside the existing schema's properties object:

{
  "properties": {
    "notionConnector": {
      "title": "Notion connector",
      "type": "string",
      "description": "Optionally write jobs to Notion through an authorized connector.",
      "resourceType": "mcpConnector",
      "nullable": true,
      "sectionCaption": "Notion output",
      "mcpServers": [
        {
          "url": "https://mcp.notion.com/mcp",
          "tools": {
            "required": [
              "notion-create-pages",
              "notion-fetch"
            ]
          }
        }
      ]
    },
    "notionDataSourceId": {
      "title": "Notion data source ID",
      "type": "string",
      "description": "Destination data source UUID or collection:// reference.",
      "editor": "textfield",
      "nullable": true
    }
  }
}

resourceType: "mcpConnector" turns the input into a connector picker in Console. After building the Actor, authorize Notion under Settings > API & Integrations, then select it in the Notion output section.

The second input, notionDataSourceId, chooses where the pages go. Keep both optional so dataset-only runs still work. Step 3 shows how to fetch the destination reference once the connection is up.

Apify Console Actor input form showing the Notion MCP connector picker under a Notion output section.

The connector picker rendered in the Actor input form, with the Notion connector selected.

The tools.required list is the part I'd skim past in someone else's article, so let me be specific about it. It isn't documentation and it isn't a hint. The proxy enforces it: it filters tools/list down to what you declared, and rejects tools/call for anything outside it.

Here's the first line my Actor logs on every run, straight from a real run log:

🔧 Tools allowed through the connector: ['notion-create-pages', 'notion-fetch']

So a jobs scraper can't rummage through your Notion. Not because I promise it won't, but because the proxy won't pass the call. You can check that yourself without reading a line of my code, which is about the only kind of security promise worth anything in a public Actor store.

One thing the allowlist doesn't do: it limits which operations the Actor can perform, not which content those operations can reach. notion-fetch isn't restricted to one database. What it can read still depends on the Notion account and permissions sitting behind the connector.

Apify Console settings page listing an authorized Notion MCP connector.

MCP connectors under Settings > API & Integrations, showing the authorized Notion connector.

Apify Console Edit connector dialog listing 28 available Notion MCP tools with read-only and idempotent annotations.

The connector's Edit dialog, showing all 28 Notion tools it could expose.

2. Connect and read the tool descriptions before you write anything

The documented snippet couldn't run on either SDK version

The first thing I did was copy the Python snippet out of Apify's connector docs. It deployed, and then:

ValueError: not enough values to unpack (expected 3, got 2)

The snippet unpacked three values from the transport:

async with streamable_http_client(...) as (read, write, _):

In the mcp 2.x SDK, streamable_http_client yields exactly two. I checked the type rather than guessing:

>>> from mcp.client.streamable_http import TransportStreams
>>> TransportStreams
tuple[ReadStream[SessionMessage | Exception], WriteStream[SessionMessage]]

Three values is the 1.x shape, where the third element was a session-ID callback. But 1.x has no function called streamable_http_client at all. It's streamablehttp_client, with no underscore between streamable and http, and it takes a headers argument rather than an injected HTTP client. I downloaded the 1.9.0 source to confirm:

# mcp 1.9.0 - mcp/client/streamable_http.py
async def streamablehttp_client(
    url: str,
    headers: dict[str, Any] | None = None,
    timeout: timedelta = timedelta(seconds=30),
    sse_read_timeout: timedelta = timedelta(seconds=60 * 5),
    terminate_on_close: bool = True,
)

So the documented snippet was a hybrid: the 2.x import name with the 1.x unpacking. It couldn't run on either version. On 2.x the import resolves and the unpacking fails. On 1.x the import fails before you get that far.

Apify has since fixed that page. It now unpacks two values, and it carries a version callout spelling out the same distinction: on mcp 1.x the transport function is named streamablehttp_client, takes headers instead of http_client, and yields a third value. If you're reading this with a working snippet in front of you, that's why.

Save the tool descriptions before you build a payload

The notion-create-pages description is 5,857 characters long. Console truncated it in my logs, so for the first few attempts I was coding against maybe a third of the contract.

The fix is one extra block, and it's the trick I'd start with on any MCP service: save the tool descriptions and input schemas into the run's key-value store, then read them properly.

Which means your first connection should do nothing but discover tools. It confirms authorization and hands you the service's real contract before you build a single payload. Each tool description explains what the tool does, and its input schema defines the arguments it accepts.

Open a session from Python

The Actor needs three values: the selected connector ID, the MCP proxy's base URL, and the Apify run token. Apify supplies the last two in every platform run.

Install these in the Actor build:

apify>=2.0.0
mcp>=2.0.0
httpx2>=2.5.0,<3
import asyncio
import os

import httpx2
from apify import Actor
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client


async def main() -> None:
    async with Actor:
        actor_input = await Actor.get_input() or {}
        connector_id = actor_input.get('notionConnector')
        proxy_url = os.environ.get('ACTOR_MCP_CONNECTOR_BASE_URL')
        token = os.environ.get('APIFY_TOKEN')
        if not connector_id or not proxy_url or not token:
            Actor.log.warning(
                '⚠️ No connector or no proxy credentials - skipping Notion sync. '
                'MCP connectors only resolve in a platform run, not a local one.'
            )
            return
        try:
            async with httpx2.AsyncClient(
                headers={'Authorization': f'Bearer {token}'},
                timeout=httpx2.Timeout(60.0),
                follow_redirects=True,
            ) as http_client:
                async with streamable_http_client(
                    f"{proxy_url.rstrip('/')}/{connector_id}",
                    http_client=http_client,
                ) as (read, write):
                    async with ClientSession(read, write) as session:
                        await session.initialize()
                        tools = (await session.list_tools()).tools
                        Actor.log.info(
                            f'🔧 Tools allowed through the connector: '
                            f'{[tool.name for tool in tools]}'
                        )
        except Exception:
            Actor.log.exception('Notion connector check failed')
            raise


if __name__ == '__main__':
    asyncio.run(main())

Then, immediately after the tools assignment and inside the same initialized session, add the block that saves everything:

payload = [
    {
        'name': tool.name,
        'description': tool.description,
        'inputSchema': tool.input_schema,
    }
    for tool in tools
]
await Actor.set_value('NOTION_TOOL_SCHEMA', payload)

Open the run's default key-value store and read NOTION_TOOL_SCHEMA. One naming detail that tripped me up: the SDK attribute is input_schema, while inputSchema above is just the JSON key I picked for the saved file. Inspect the allowed arguments for notion-fetch and notion-create-pages before going further, because the parent ID and the property formats you need in step 3 are both buried in there.

Push to a dev build tag, not latest. I have paying users on latest, and I wasn't going to test a new network call in front of them. There's a second reason to work this way: ACTOR_MCP_CONNECTOR_BASE_URL only exists in a platform run, so apify run on your laptop can't reach a connector at all. My whole loop was push to a dev tag, run on the platform, read the log.

3. Prepare the Notion destination and map its fields

This is the step that cost me the most time, and the worst failure in it was the silent one.

A batch of pages went into Notion with nothing but titles. No company, no salary, no equity. The log cheerfully reported that it was sending all columns. It wasn't lying, exactly: my schema parser had returned an empty dict, and the row builder read {} as "this database has a title column and nothing else", so it filtered every other field out on the way through.

Two things have to be right before you write anything: the destination reference, and the column map.

Get the data source reference

My first mistake here was simpler. I sent the database ID, because that's what the Notion URL hands you, and the tool rejected it.

A Notion database contains one or more data sources, and the write needs the ID of the source that should receive the jobs, not the database itself. A database URL on its own isn't the destination value either. Fetching the database is what makes the distinction obvious.

Set database_url to the URL copied from your Notion database. Add this after session.initialize() in the connection example, at the same indentation as the other statements inside that session:

result = await session.call_tool(
    'notion-fetch', arguments={'id': database_url}
)
if result.is_error:
    raise RuntimeError('Notion could not fetch the database')
await Actor.set_value(
    'NOTION_DATABASE',
    result.model_dump(mode='json', by_alias=True),
)

Open NOTION_DATABASE in the run's key-value store and find the collection:// reference for the data source you want. Put that reference, or its bare UUID, into notionDataSourceId. The sync strips collection:// before sending parent: {"data_source_id": source_id}.

Read the schema before building rows

Back to the empty dict. But there's a reason I was reading the schema at all, and it wasn't the bug.

I had it working against my own database and I nearly shipped that. But it needed to work for my users' databases too. My database has a column called Salary Min. Theirs might call it Salary (min), or have no salary column at all. The title column is usually Name, but Notion lets you rename it to anything. Since one unknown column kills an entire batch, an Actor that assumes my layout works for exactly one person.

So the Actor reads the target before writing to it. That's what the second declared tool, notion-fetch, is for.

In src/notion_sync.py, _fetch_data_source_schema calls notion-fetch for source_id and passes the response to parse_data_source_schema. Those helpers, and logger, are defined in that file. The sync uses their result like this:

schema = await _fetch_data_source_schema(session, source_id) or None
if schema:
    logger.info(f'🗂️ Target columns: {sorted(schema)}')
else:
    logger.warning('⚠️ Could not read the data source schema - sending all columns')

That or None is the fix for the silent failure. An empty dict is falsy, but it isn't None, and the row builder treated {} as a real schema with no optional columns. Collapsing it to None forces a failed or empty lookup down the fallback path: send the expected fields and let Notion report the mismatch out loud. You can still end up with a rejected batch, but you'll know about it instead of finding fifty half-empty pages later. And either way the complete scraped data is already sitting in the Apify dataset.

The parsing itself has a wrinkle. notion-fetch doesn't return JSON. It returns a document wrapped in a JSON envelope: some prose, a <data-source-state> block, and a SQLite CREATE TABLE rendering of the same schema. The parser has to unwrap the envelope before it can read the schema at all. This excerpt handles the outer layer; the repository's parse_data_source_schema handles the remaining formats:

import json

try:
    envelope = json.loads(payload)
    if isinstance(envelope, dict) and isinstance(envelope.get('text'), str):
        payload = envelope['text']
except json.JSONDecodeError:
    pass

Then the row builder sends the job title to the column with type title and filters the other fields against the discovered names. This excerpt uses _title_column and the candidates dictionary from build_job_rows in the same file:

title_key = _title_column(schema) if schema else 'Name'
...
for column, value in candidates.items():
    if value is None:
        continue
    if schema is not None and column not in schema:
        continue
    properties[column] = value

The nice side effect: point the Actor at a brand-new Notion database with only a title column and it still works. It writes titles. Add a Salary Min column and the next run fills it in, with no code change.

Use the MCP tool's property formats

One more trap, and this one produced the least helpful error message of the entire build:

{"code":"validation_error","status":400,
 "message":"Properties {propertyKeys} not found in the data source.\n
            Date property {propertyKeys} not found in the data source.\n
            All editable property keys: {editableProperties}."}

Those are literal {propertyKeys} placeholders. The template never got interpolated, so the error tells you that some property is wrong while withholding which one and what the valid names are. Notion knows both. And one unknown column rejected all 20 pages in the batch, because the offending page doesn't fail on its own.

Reading the saved tool description got me there faster than guessing from that message ever would have. It also handed me the encoding rules I'd otherwise have found by trial and error. Properties take scalar values rather than the Notion REST API's nested property objects. Send numbers as numbers. Checkboxes are the literal strings __YES__ and __NO__. And date properties split across prefixed keys, so a column called "Scraped At" is written like this, where properties is the row's property dictionary and scraped_at is its date value:

properties['date:Scraped At:start'] = scraped_at

Not properties['Scraped At']. That one would have taken me an hour on my own.

Match column names exactly and give them compatible types. A Number column named Salary Min will take the scraper's salary value; Salary (min) gets skipped. The Actor discovers the title column, but it won't translate other names for you or validate every column type.

4. Save the scrape, then write pages in batches

My first connection attempt broke, and the run still finished with exit code 0 and a complete dataset.

That wasn't luck, it's the entire reason for the ordering in this step. Save the scrape to the dataset first, then sync, and keep the sync inside its own try/except. A connector failure should never throw away work the scraper has already finished, and on a per-result Actor that's work the user has already paid for.

Here's the call site in src/main.py, after the results are saved:

# Dataset first - this must succeed before anything touches Notion
await Actor.push_data(jobs)
notion_connector = actor_input.get('notionConnector')
notion_data_source_id = actor_input.get('notionDataSourceId')
if notion_connector and notion_data_source_id:
    try:
        created = await sync_jobs_to_notion(
            jobs,
            connector_id=notion_connector,
            data_source_id=notion_data_source_id,
        )
        Actor.log.info(f'✅ Wrote {created} jobs to Notion')
    except Exception:
        Actor.log.exception('Notion sync failed - dataset is unaffected')
else:
    Actor.log.info('No Notion connector selected - dataset output only')

The repository has the surrounding Actor lifecycle and job collection. The two guards that matter are both here: only call the sync when notionConnector and notionDataSourceId are both supplied, and never let it raise past the try.

Check tool responses as well as exceptions

A successful network request can still carry a failed tool result. In the Python SDK the flag is is_error, snake case, not the isError that most existing MCP writing shows, because most of that writing is TypeScript. Exceptions cover transport and protocol failures; is_error covers everything else, and you have to check it after every call. Miss it and every write looks like it worked while nothing appears in Notion.

Define this helper in src/notion_sync.py before the batch-writing function:

from typing import Any, Optional


def _tool_error_text(result: Any) -> Optional[str]:
    if not getattr(result, 'is_error', False):
        return None
    for block in getattr(result, 'content', None) or []:
        text = getattr(block, 'text', None)
        if text:
            return text
    return 'unknown error'

Once the session is initialized and the rows are built, call notion-create-pages with the data source ID as parent. This excerpt from sync_jobs_to_notion uses _chunk and PAGE_BATCH_SIZE (20) from the same file; CREATE_PAGES_TOOL is notion-create-pages, and created starts at zero:

for batch_no, batch in enumerate(_chunk(rows, PAGE_BATCH_SIZE), 1):
    response = await session.call_tool(
        CREATE_PAGES_TOOL,
        arguments={
            'parent': {'data_source_id': source_id},
            'pages': batch,
        },
    )
    error_text = _tool_error_text(response)
    if error_text:
        logger.error(f'Notion batch {batch_no} failed: {error_text}')
        continue
    created += len(batch)

Log the failed batches and count only the successful ones. And since the run can exit 0 with a broken sync, check the sync logs, not just the overall run status.

Notion database of Y Combinator jobs with company, YC batch, salary, and equity columns filled in by the Apify Actor.

The Notion database after a run, populated with YC jobs.

Start with a small run and check a few pages against the dataset: job title, company, salary, destination. If your database only has a title column you'll get title-only pages, which is the same symptom as the empty-schema bug in this step with a far more boring cause. Add the optional columns before you expect to see values in them.

5. Verify the result and measure the overhead

Two seconds and about 5% more compute. That's what the Notion write cost me.

I ran the Actor over 20 companies at 512 MB, with and without the connector:

Runtime Compute units Jobs Notion pages
Without connector 38.7 s 0.00537 50 0
With connector 40.7 s 0.00566 50 50

Apify run log showing the MCP connector tool list, the discovered Notion columns, and fifty jobs written in three batches.

The Apify run log showing tool discovery, the target columns, and the batch writes.

Because this Actor charges per result, that extra compute comes out of my margin. It doesn't raise the per-result price my users pay. At 5% I'll take that trade happily, but it's worth knowing which side of the ledger the overhead lands on before you ship a connector on a per-result Actor.

The 50 pages went out in three notion-create-pages calls, batched 20, 20, and 10. Those were the writes only; tool discovery and schema fetching were additional requests. One call per job would have meant 50 round trips and a much less pleasant table. It's a small comparison, so treat it as the overhead I measured rather than a guarantee for larger runs or other services.

The scrape output had the same shape in both runs: 20 companies, 50 jobs, 28 founders. The connector is optional. If nobody selects one, the Actor makes no MCP calls at all, and my existing users see exactly what they saw last month.

Before you schedule it, deal with duplicates

One thing I haven't solved yet. Every run appends new pages, so a daily schedule will re-add jobs that are already in the database.

There's no deduplication in the current implementation. If you want a persistent job tracker rather than a snapshot, comparing job URLs before writing is the next change to make, and I'd make it before setting up any schedule.

What I'd do differently

I'd read from Notion, not just write to it. Right now the filters live in the Actor input. It'd be better to keep a Notion database of the batches, industries, and locations I care about, have the Actor read its own task list at the start of a run, and write results back. A loop instead of a one-way pipe. The connector already permits it; I just haven't built it.

I'd also like to try two connectors on one Actor: Notion for the archive, Slack for a message when a job matches a saved search. The input schema accepts an array of connector IDs, with one MCP session per connector.

Reuse the Apify MCP connector pattern in another Actor

For a different destination, keep the same order: declare the tools, connect and save their descriptions, inspect the destination, map your results, and write only once the dataset is safe.

Another service will have its own connector, its own tool names, and its own payload format. Which is exactly why the discovery step in step 2 is worth doing first. It hands you those requirements up front instead of making you reverse-engineer them from error messages with {propertyKeys} in them.

FAQ

Can I add a connector to any existing Actor?

The Actor has to declare the connector input and contain code that calls it. For this scraper, select Notion and supply notionDataSourceId. Adding a connector in your account settings alone doesn't make an unrelated Actor write to Notion.

Can the Actor see my Notion token?

No. It authenticates to the Apify MCP proxy with its own run token, and the proxy adds the Notion credential server-side. The connector session ends with the run.

Can I restrict what the Actor can do?

Yes. Access has to satisfy the service's authorization permissions, any connector-level tool allowlist, and the Actor's mcpServers[].tools declaration. A tool allowlist limits operations, while the upstream service controls which content those operations can reach.

Can I develop this locally?

Not really. ACTOR_MCP_CONNECTOR_BASE_URL only exists in a platform run. I test row building and response parsing locally, then use a dev build to verify authorization and real tool calls.

Which SDK version do I need?

mcp 2.x for Python, which pulls httpx2 rather than httpx. The 1.x API differs in both the transport function's name and its signature.

Is this the Apify MCP server?

No. Here the Actor calls Notion through an outbound MCP connector. The Apify MCP server is the other direction: it lets external AI clients call Actors.


Code: the connector implementation is in src/notion_sync.py.

The Actor: YC Jobs Scraper on Apify Store.

Further reading

来源:Google AI:DEV 作者专属(RSS) · dev.to