平台团队如何评估 AI 生成的申请信:从人工标准到 LLM 评审校准
how we evaluate our AI Gen message for apply message
作者团队为平台上的服务者构建了 LLM 自动撰写申请信的功能,并复盘了如何评估生成质量。第一版评估工具因标准出自个人想象、未与人工标注校准,fabrication 评审的 precision 为 0,导致新旧 prompt 对比结果不可信。
how we evaluate our AI cover letter
we built a feature: when a pro (service provider on our platform) gets a request from a client, LLM writes the first apply message for the pro.
there are 2 routes. fullAi writes the whole letter, insert adds one AI paragraph into the pro's own template.
after release, the hard question was: is the generated letter good or not ?
this post is how we try to answer it.
1. why we need evaluation
we can't tell good or bad by feeling.
one real example: a letter for a water leak request at an apartment asked the client "where on the upper floor is the leak coming from ?"
the upper floor is someone else's home, so the client has no way to answer it.
at that time I ran all of our rule checks (contact info, placeholders, price source...), none of them caught it.
I found it by reading.
and the same request generated 3 letters, only 1 of them asked this. which is the randomness of the generation itself.
so we need 2 things:
- a golden standard for what is a good letter. at the beginning only humans can define it.
- an automated evaluation system. human review is expensive, reading tens of letters every time we change the prompt is not sustainable.
the order matters. humans define the standard first, then let the machine learn it. and the machine judge also needs to be verified.
I learned this after doing it in the wrong order once.
2. round 0: the mistakes I made
one month before release, I built a first version of the eval tool.
looking back, almost every step had a problem.
what it looked like:
- a checklist of 10 rules. 6 are code checks: link, placeholder left in the text, contact info, length, prompt leak, refusal phrases. 4 go to an LLM judge: can the client answer it, fabrication, commitment, and saying "we can do it" only at the category level.
- one prompt per judge rule. quote the evidence first, then return true / false. if not sure, it fails.
- the data is letters generated from real requests at our QA environment, plus synthetic violation samples: commitment 20, fabrication 20, injection 24, and 20 clean ones.
mistake 1: the standard came from my head, not from humans
I wrote the rules from my imagination, there were no human labels behind them.
later when I calibrated with human labels, the fabrication judge had precision 0. none of the issues it flagged was a real issue.
the answerability judge was 0.33, so 2 out of 3 flags were wrong.
mistake 2: comparing prompts with an uncalibrated judge
end of August I changed the prompt to stop saying "we can do it" only at the category level.
for the same 51 requests, the old and new prompt generated once each, then the judge scored them:
| judge | old prompt | new prompt |
|---|---|---|
| "we can do it" at category level | 28/51 (55%) | 15/51 (29%) |
| commitment | 10/51 (20%) | 5/51 (10%) |
| fabrication | 12/51 (24%) | 14/51 (27%) |
| client can't answer | 8/51 (16%) | 6/51 (12%) |
the first 2 rows look great, both cut in half.
but fabrication went up from 24% to 27%, and this judge has precision 0. so this row tells nothing.
same for the other rows. every number in the table comes from the judge,
and at that time I could not answer whether the judge itself was right.
so the comparison can only be as reliable as the judge behind it.
mistake 3: "not sure" counted as fail
this way the violation rate mixed in the judge's own uncertainty.
we can't tell if the letter has a problem, or the judge is just not sure.
the later version uses 4 values, and "not sure" is recorded separately as uncertain (see section 4).
mistake 4: one checklist for 3 different things, judged by the literal text
- 3 kinds of things in one list: what the generation must not do (e.g. no contact info), how the whole letter reads (can the client answer it), and a fixed length limit.
- no split for the template. contact info, prices and repeated questions from the pro's own template were also counted as the AI's problem.
- judging by the literal text gives false flags. e.g. "we know this area well" was treated by the fabrication judge as a fact that needs a source.
we didn't throw the checklist away, it's kept as a reference.
placeholder left, prompt leak and refusal phrases later became candidates for the basic defect check.
so round 1 started from scratch: find the people, define the standard, then let the machine learn it.
3. find the right people: domain experts
engineers can't judge if a cover letter is good, because we don't know what makes a client want to reply.
we asked CS and Sales to read real samples and write feedback. a lot of it is something engineers won't think of, for example:
- "repeating the request at the beginning of the letter is not needed."
- "instead of asking the client's current weight, ask how much they want to lose, or which body part they want to work on."
this feedback became prompt changes directly, and also the material for the eval standard.
4. collect small samples: good, bad, and not sure
human time is limited, so we only look at 20-30 letters each time.
- 1st batch: from the first 500 letters after release, stratified random sampling by route, 15 each, 30 in total.
- 2nd batch: another 20 (10 for each route), no overlap with the 1st batch, kept only for validation.
we don't give each letter one overall score. we label per dimension, with one of 4 values:
- acceptable
- needs improvement
- not applicable
- uncertain (not enough evidence, or the boundary of the standard is not decided yet, so don't force a pass or fail)
one trap: no issue marked doesn't mean pass.
a dimension the reviewer didn't mention is "not checked", not "acceptable".
5. when people disagree, discuss and write down the standard
in the 1st batch of 30, the human overall rating was good 24, ok 6, bad 0.
most letters have no big problem, the problems are in the details.
so a simple good / bad is not useful, we need to split it into concrete dimensions.
the 10-rule checklist from round 0 is not used by default anymore, the reason is mistake 4 in section 2.
after going through each letter together with the original request, we defined 5 dimensions:
| dimension | key in code | what it checks |
|---|---|---|
| core need | core_need |
when it's still unclear what the client wants done, ask that first, before the work details |
| reply burden | reply_burden |
can the client answer easily ? don't ask for technical categories or a full set of documents first |
| alternative fit | alternative_fit |
"if not sure, send a photo": can the photo actually answer the original question ? |
| assembly | assembly |
does it ask again what's already in the request form, and is the paragraph order natural |
| intent | intent |
does it respond to the real purpose in the client's comment |
each dimension is judged separately, no total score.
a total score mixes "one serious problem" and "a few small issues" into the same number.
6. let the LLM judge by the standard
we didn't automate all 5 dimensions at once.
the rule is to only automate dimensions that actually showed up, and whose boundary is already discussed.
we started with 2: core_need and reply_burden. the other 3 are still judged by humans.
what the judge outputs
for each letter, the judge returns 3 verdicts on 2 axes:
- business quality (
business):core_need,reply_burden. it looks at the whole letter, not only the AI paragraph. - prompt compliance (
prompt_compliance): only checks if the AI paragraph breaks the generation instructions of that time. e.g. facts with no source, treating a reference price as the quote for this request, promising it can take the job, leaking the instructions, breaking the format rules of the route.
the 2 axes are judged independently, no total score.
a letter that follows the generation instructions perfectly can still need improvement on business quality.
so splitting them tells us which one to fix: the model not following the prompt, or the prompt itself.
each verdict has 5 fixed fields:
| field | meaning |
|---|---|
verdict |
acceptable / needs_improvement / not_applicable / uncertain, the same 4 values as the human labels |
output_quote |
exact quote of the problem from the letter |
input_quote |
exact quote of the evidence from the generation input |
reason |
the reason for the verdict |
responsibility |
which layer the problem comes from, see below |
responsibility is the most useful field for us.
one letter is assembled from several parts, so the same problem can need a fix at a totally different place:
-
generated_text: the AI paragraph. fix the generation prompt -
template_or_assembly: the pro's own template, or lines added during assembly. not the AI's problem -
source_context: the input itself is missing info or has a conflict -
unclear: can't tell -
none: no problem
one real output (a rain gutter repair request, it comes back in section 8).
the letter and the judge output are in Japanese, so the values below are translated:
{
"verdict": "needs_improvement",
"output_quote": "If you don't mind, around which part did you notice the problem: the eaves, the downspout, or the catch basin?",
"input_quote": "Request category: rain gutter repair / construction contractor",
"reason": "The request form doesn't specify what symptom the rain gutter has or what work the client wants. The letter should first ask about the situation, such as a leak, a clog or a detached part, or the work they want. Instead it assumes there is a problem and only asks where it is.",
"responsibility": "generated_text"
}
hard checks in the code
asking "please quote the original text" in the prompt is not enough.
the model may rewrite the quote or miss a field, so after the output comes back, the code checks it again:
- structured output with a strict schema. one more or one less dimension or field, this call fails.
- quotes must be exact substrings.
output_quotemust appear in the generated text or the whole letter,input_quotemust appear in the generation input. if not, this call fails and we don't take the verdict. -
needs_improvementmust come with a problem quote, and the reason can't be empty. - the whole config is frozen into 1 hash: the full rubric, the model, the output schema, the call parameters, and the judge code itself. if any of them changes without a new freeze, the tool refuses to run. so every score can be traced back to the judge version that gave it.
- no auto retry, every letter runs exactly 2 rounds, for checking if the judge is stable by itself.
7. round 1: find the disagreements, check whose problem it is
1st batch, 30 letters, 2 rounds each, 60 calls in total.
first, is the judge stable by itself ?
same letter, 2 runs, different verdicts: core_need 2/30, reply_burden 6/30.
where the judge doesn't even agree with itself, no need to compare it with humans yet.
then, the disagreements with humans. we picked 23 verdicts and reviewed them one by one:
the ones that changed between rounds, the ones flagged as a problem in both rounds, and spot checks of the ones that passed in both rounds.
for reply_burden, round 1 matched only 6 out of 11, almost half didn't match.
after going through them, the disagreements came from 3 places:
- the judge is too strict. e.g. it insists "must ask the specific goal first". a client wants a personal trainer and picked "health improvement" as the goal, and the letter first asks how many times a week they want to train. the judge said "should ask what they want to improve first" in both rounds. the human said it's fine, the specific goal can be asked at the trial lesson.
- the standard is not clear. e.g. whether a light re-confirmation means the letter doesn't understand the need, the old standard didn't say.
- the human label itself needs a fix. some labels had only a verdict with no reason, some changed after the review. so a disagreement is not always the LLM's fault.
for 1 and 2, we only added 3 general paragraphs to the standard:
- when the target and purpose of the request are already clear, the first letter doesn't need to confirm every detail. it can be asked at the trial or the site visit.
- a light re-confirmation from the template doesn't count as not understanding the core need.
- don't judge by the number of questions or photos mechanically. tell the difference between "could be shorter" and "the burden is really too much".
one rule: no exceptions for single cases.
if we add special cases just to match every dev sample, the judge overfits.
8. validate on new samples
using the samples we tuned the standard on to prove the standard works is like grading your own exam.
so validation uses the 2nd batch, the 20 letters that were not used for tuning.
the order also matters: humans label blind first, then we run the model.
so the human is not influenced by the model's verdict.
we set a working threshold before running:
- each dimension, each round, at least 18/20 agree with the human (disagreement <= 10%).
- verdict changes between the 2 rounds no more than 1/20.
how we count the disagreement:
- false flag: the human says acceptable, the judge says needs improvement.
- missed issue: the human says needs improvement, the judge says acceptable.
- only items where both sides give acceptable or needs improvement go into the denominator. if either side says not applicable or uncertain, or the human label is not confirmed yet, it's counted separately, neither right nor wrong.
- by default the tool only compares items where the 2 rounds agree, the ones that changed are counted as unstable. the table below is per round.
the result:
| dimension | round 1 agrees with human | round 2 agrees with human | same verdict in 2 rounds | false flags (round 1 / 2) | missed issues (round 1 / 2) |
|---|---|---|---|---|---|
| core need | 16/20 | 15/20 | 19/20 | 4 / 5 | 0 / 0 |
| reply burden | 16/20 | 14/20 | 18/20 | 2 / 4 | 2 / 2 |
all the core_need disagreements are false flags, the judge is stricter than the human and tends to require "ask the goal first".
reply_burden has errors in both directions.
the disagreement rate is 20-30%.
better than the dev stage, but it hasn't reached the threshold yet: core_need only passed the stability one, reply_burden passed neither.
one more limit: in these 20, only 1 letter was labeled by the human as having a core_need problem.
with only 1 negative case, it can't prove the judge is able to catch problems.
so for now, this judge alone can't tell us "the new prompt is better than the old one".
the next step is to review these disagreements and iterate again.
a typical disagreement: the client only wrote "rain gutter repair", and the letter asks "is the problem around the eaves, the downspout, or the catch basin ?"
the judge said "should ask what the symptom is first" in both rounds, the human said it's acceptable.
this is exactly the kind of case we need to bring back to the original request and discuss.
9. after reaching the threshold: release small
after the judge is stable, then we use it to change the production prompt.
if there is no urgent release date, start small:
- shadow mode: the new prompt generates in production in parallel, but doesn't send to clients.
- use these real production samples for a 2nd round of human labeling.
- run the LLM judge again and check the disagreement rate.
- if the agreement is good enough, roll out gradually. if not, go back and review the disagreements one by one like section 7.
the goal is to keep improving the production prompt with confidence.
lessons
- calibrate the judge first, then use it to compare prompts. with an uncalibrated judge, you can't tell if the letter changed or the judge is just flagging randomly.
- no total score. judge per dimension, so the problem is visible.
- not labeled doesn't mean pass. record "not checked" and "acceptable" separately.
- a disagreement is not always the LLM's fault. human labels get fixed, and the standard gets completed too.
- the validation set can't be used to tune the standard, otherwise it's grading your own exam.
- run every item 2 rounds. where the judge isn't stable by itself, don't compare it with humans yet.
- the threshold is a working threshold, not a statistical proof. 20 samples can only say so much, say it clearly.
来源:Google AI:DEV 作者专属(RSS) · dev.to