如何用 Chrome 扩展读取任意电商产品页:JSON-LD 优先,选择器兜底
Reading any marketplace product page from a Chrome extension: JSON-LD first, selectors second
一款名为 AI Product Page Optimization 的 Chrome 扩展,通过先解析页面中的 application/ld+json 结构化数据、再按平台加载限定范围的 DOM 选择器,读取 Amazon、Shopify、eBay、Etsy 产品页并约 30 秒返回标题、描述与图片建议。
If you have ever tried to build a tool that "reads" a product page — an Amazon listing, a Shopify storefront, an eBay item, an Etsy shop page — you already know the problem: every marketplace renders the same conceptual data (title, bullets, description, images, price) in completely different DOM shapes.
We ran into this while building AI Product Page Optimization, a Chrome extension that reads the product page you already have open and returns marketplace-ready title, description, and image suggestions in about 30 seconds. This post is the technical breakdown of how the reading step works. (This article is disclosed as AI-assisted; the extraction design and constraints below are from our actual implementation.)
Start with JSON-LD, always
Most marketplaces embed structured data in the page. Before touching a single DOM selector, we look for application/ld+json blocks and parse them for Product-shaped objects:
const ldjson = Array.from(
document.querySelectorAll('script[type="application/ld+json"]')
).map((s) => {
try { return JSON.parse(s.textContent); } catch { return null; }
}).filter(Boolean);
Why JSON-LD first:
- It survives layout changes. Amazon can reshuffle its DOM next week; the
Productschema block rarely moves. - It normalizes the shape. A title is a
name, images are animagearray, and offers carry price/currency regardless of which marketplace you are on. - It handles encoding edge cases (HTML entities, unicode in internationalized catalogs) better than regex-over-HTML ever will.
But JSON-LD is never the whole story. Marketplaces routinely put richer or more current data in the DOM than in their structured data: bullet points that are not in any schema field, variant-specific copy, A/B-tested titles.
Selectors as a scoped fallback
After JSON-LD, we apply platform-specific DOM selectors to fill the gaps — and the key word is scoped: we detect the marketplace from the URL first, then load only that platform's selector set.
The four surfaces we read today:
- Amazon — deep, nested DOM with frequent class-name churn; stable attributes and landmark patterns beat styling hooks
- Shopify (including custom domains) — the most uniform target; Shopify themes share enough structure that extraction is reliable even on stores with custom themes
- eBay — item-specific layouts where the description block is seller-authored HTML, which changes the parsing strategy entirely
- Etsy — listing copy with a distinct handmade/creative tone that also matters downstream for generation
Constraints matter more than extraction
Reading the page is half the job. The other half is knowing what each marketplace will accept when you write the optimized copy back:
- Amazon truncates aggressively on mobile — character limits are hard product constraints, not suggestions
- Shopify storefronts often carry thin manufacturer text that needs brand-voice cleanup rather than keyword stuffing
- eBay titles have a hard length cap, and condition-forward descriptions convert better
- Etsy rewards handmade-tone phrasing with stronger search phrases — generic copy reads as dropshipping
One more detail that surprised us: output language follows the detected page language. Cross-border catalogs mix English with other locales, and returning English suggestions for a Japanese listing page is worse than useless — it breaks the seller's workflow.
MV3 practicalities
Two things worth knowing if you build this on Manifest V3:
- The extraction needs to run in the page context (JSON-LD and DOM both), so your content script is the workhorse; the service worker only orchestrates.
- Keep the extension stateless about page data — read the live page on demand, send only what the user approves onward. It keeps you compatible with platform policies and keeps the data footprint small.
Review before publish
We deliberately never auto-paste anything. Extracted data and generated suggestions land on a web dashboard with a before/after comparison, history, and optional share links — the seller copies and publishes only what they approve. Optimization is credit-based (text 1 credit, AI-enhanced image 2, keeping an original image is free; a typical product page uses about 5 credits, and new accounts start with 5 free credits).
If you are building anything that reads product pages — price trackers, feed tools, listing optimizers — the pattern is the same: structured data first, scoped selectors second, platform constraints third, and a human review step before anything goes back to the marketplace.
Try the extension: aiproductpageoptimization.com
来源:Google AI:DEV 作者专属(RSS) · dev.to