引用Dr. Cheems Wang 🏡@AlbertW24045555
You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
Picking the best few-shot demos is a slow System-2 search: combinatorial, often with repeated LLM calls.
We made it System 1. ⚡
Jev-LDE, a 1.7B editor, glances at the retrieved demos and makes ONE edit. The LLM answers once.
Avg 1-shot acc 81.2 → 88.1
You Only Edit Once 🧵
Animated walkthrough (illustrative example). Query: "How far is it from Denver to Aspen?" Semantic TopK retrieves three look-alike demos: "Where is Aspen, Colorado?" (Location), "What state is Denver in?" (Location), "Who founded Denver?" (Person). Jev-LDE, a 1.7B System-1 editor, flags the first as same topic but wrong answer type and outputs one action: Replace S1 with candidate C1, "How far is Boston from NYC?" (Number). The frozen target LLM then answers "Number", which is correct; without the edit it answers "Location". End card: average 1-shot accuracy 81.2 to 88.1 across 3 benchmarks and 4 target LLMs, best or tied-best in 44 of 48 settings, +11% wall time.