如何构建基于照片的卡牌评级系统:让 LLM 只描述不打分的计算机视觉方案
Building a Photo-Based Card Grader: Computer Vision Where the LLM Doesn't Pick the Score
CardGrade 用手机照片预评卡牌,其 CGI Vision AI 引擎约 60 秒返回 PSA、BGS、CGC 估算评级及置信度。架构上由独立 Python/FastAPI 服务用 OpenCV 负责像素处理,视觉模型只描述缺陷(none/minor/moderate/severe),由确定性评分规则换算分数,避免 AI 直接打分。
Trading-card grading looks like an image-classification problem until you try to build it. A grade is driven by several physical properties at once: how centered the art is, how sharp the corners are, whether an edge is whitening, whether the surface has a fold. Each one lives at a different scale in a phone photo.
This is a write-up of how we structured the pipeline behind CardGrade, an app that pre-grades cards from phone photos. Its engine, CGI Vision AI, returns estimated PSA, BGS and CGC grades with confidence scores in about 60 seconds. It's an architecture post, not a benchmark post. I'm not quoting an accuracy number, because a photo-based result is an estimate, and I'd rather say so than dress it up.
The stack, briefly
- Web: Next.js, React and TypeScript, styled with Tailwind
- Data: PostgreSQL with Drizzle ORM
-
Mobile: React Native with Expo, using NativeWind for Tailwind-style classes and
react-native-vision-camerafor capture - CV service: a separate Python/FastAPI service running OpenCV
- Packaging: Docker containers for the web app and the CV service
Why one service owns the pixels
The grading request is small. The web app stores the full-resolution originals and sends the CV service a payload with a grading ID, front and back image URLs, and a webhook URL. The service replies 202, does the work, and calls the webhook with structured results.
The reasoning: the CV service is the only place where the original pixels, the detected card border, and every consumer of the crops meet in a single request. If the web app also cropped images, we'd maintain two implementations of the same geometry, and they would drift.
Step 1: find the card, then flatten it
The first model is a border detector. It finds the card's outer edge and the inner edge of the printed border. Those two quadrilaterals drive everything downstream.
From them we compute a perspective transform and warp each side onto a flat canvas. Small details bite here:
-
Corner ordering must be unambiguous. Top-left minimizes
x + y, top-right maximizesx − y, bottom-right maximizesx + y, bottom-left minimizesx − y. Rotated quads are a test case. - Warp parameters matter. Bilinear interpolation with replicated borders, so we don't invent edge pixels.
- Keep a margin. The canvas includes a little background around the card for edge contrast.
Once the card is flat and axis-aligned, zones become slices of one canvas per side: a fixed fraction of the card width for each corner, a thin band for each edge, and the inner art area for surface. Four corners and four edges on two sides gives the 16 inspection zones CGI Vision AI reports on. Centering and surface are assessed alongside them.
Step 2: don't throw away resolution
An early version of the pipeline upscaled small uploads to roughly 4000px on the long edge before doing anything else. The border model was tuned around that scale, and many absolute-pixel thresholds were calibrated to it.
The catch is that upscaling is interpolation, and interpolation is a low-pass filter. Fine scratches are high-frequency detail, and they're gone before a crop is cut. Corner and edge geometry is low-frequency, so it survived. Surface defects didn't.
Removing the upscale outright would change three things at once: the border model's input distribution, the polygon coordinate space, and every absolute-pixel constant. So the V2 cropper uses a two-source model:
- Keep the upscaled image only for border geometry.
- Map the detected polygons back into original-pixel space.
- Cut every defect crop from the untouched original.
The V2 cropper is built around this. It is a Python port of the crop geometry with a versioned geometry registry, checked against the original TypeScript implementation on real-photo fixtures to sub-pixel parity. Zone crops are written under a -v2 naming scheme with a geometry sidecar, so every crop can be traced to the exact geometry that produced it.
Step 3: classical CV for what can be measured
Anything measurable gets measured with OpenCV instead of guessed:
- Centering comes from border geometry. Average the left and right border widths and the top and bottom widths, express them as percentages (say 48/52), and convert using standard card dimensions of 63.5 × 88.9 mm.
- Corners and edges get per-zone metrics such as fray, fill, angle, whitening and rounding for corners, and wear measures for edges.
Step 4: a vision model that describes, never scores
For defects that are hard to hand-engineer, a vision model examines each zone crop and describes what it sees, with a severity of none, minor, moderate or severe. The division of labor is strict: the AI never picks scores, it only describes. A deterministic rubric converts those observations and the CV metrics into numbers.
That split means you can trace why a card scored what it did, and changing a scoring rule doesn't mean re-prompting anything. Caps are explicit too: a creasing finding caps the surface score by severity.
Each subgrade uses the weakest zone: corners is the minimum of the four corners, edges the minimum of the four edges, with no averaging. The overall grade is a weighted blend of the four subgrades, rounded to the nearest half point. Centering carries the least weight and surface the most. If centering can't be measured, it is dropped and the remaining weight is redistributed rather than filling in a made-up number. On top of the blend sit weakest-link caps, so one badly damaged pillar can't be averaged away.
Those weights are our model of the process, not anything PSA, BGS or CGC publish. Results are estimates, not official grades, and CardGrade is independent of all three.
The hard part: surface
Surface is where a single phone photo is weakest. Some professional systems build a surface topology map from multiple lighting angles. A phone photo gives you one lighting condition.
Corners and edges have measurement channels. Creases are harder, because a fold is a three-dimensional event and a photo is flat. A model's opinion alone can go wrong in both directions: a real fold can be missed, and a print line can be read as a crease. So crease handling is layered:
- The vision model describes creasing as one of nine surface families for each face, and it may answer "uncertain" or "unobservable" instead of being forced into "none".
- A separate surface-defect model can corroborate a crease. A lone moderate-or-worse crease call on an otherwise clean card is checked against it before it is allowed to cost the card.
- The rubric caps the overall grade by crease severity, because a crease binds the whole card, not just the surface subgrade.
- Geometry-based crease analysis is part of the toolkit too. It treats a wrinkle (visible on one side) differently from a crease (apparent on both sides), uses cross-side registration as evidence for the more severe class, and treats print lines as the main false-positive class, since a fold changes the surface normal and ink does not.
Mobile notes
The member-facing product is native iOS and Android on React Native and Expo. JS-only changes ship over the air with EAS Update. The runtime version is tied to the app version, so an OTA only reaches matching binaries. When a change alters the client/server contract, keep the server backward compatible with clients that may never update.
What I'd tell someone starting out
- Put image processing in one place and keep the app a thin client.
- Preserve original pixels for anything high-frequency. Upscale only for what needs it.
- Measure what's measurable, and let a model describe only what isn't.
- Keep the model out of the final number. A rubric you can read beats a score you can't explain.
- Be honest about the ceiling. A photo can't see everything a grader sees under magnification, so the output is an estimate.
If you want to see the result, it's at cardgrade.io. Questions about the pipeline are welcome in the comments.
来源:Google AI:DEV 作者专属(RSS) · dev.to