跳转到内容
Rossovia方法仓库

Model Evaluation

用重复的真实任务证据形成有边界的模型执行能力画像。

源文件边界

本页为人类读者重排 skills/model-evaluation/SKILL.md 中的关键判断,不替代可安装的完整 prompt;事实、修改权和未展示的运行说明仍属于源文件。

Terminal window
npx skills add lidessen/rossovia --skill model-evaluation

安装后直接向 agent 描述目标;只有在需要强制选择时才显式点名 skill 或内部 operation。

Own one recurring judgment: what traceable task evidence justifies a bounded capability claim about an execution profile strongly enough to inform later allocation?

Evaluate an execution profile, not a model label. Its identity includes the model and provider or plan, harness and version, prompt/skill/context policy, tools and permissions, and execution policy. A changed member creates a new or revised profile unless evidence shows that the difference is immaterial.

The result is a versioned capability claim with evidence, scope, uncertainty, and a reopening observation. It is not a universal intelligence score.

Primary: P13 Supporting: P02, P05, P08

  1. Bound the claim. State the task shape, risk, and decision the profile may influence. “Better model” is not a bounded claim; “reliably reviews local TypeScript changes before human confirmation” may be.
  2. Freeze the compared conditions. Record every identity member and keep task packet, fixture revision, skills, tools, permissions, completion contract, and execution policy matched unless that member is the variable under evaluation. Make the declared effective inference policy explicit: thinking or reasoning mode and effort, temperature, streaming, context/history handling, step and duration boundaries, and completion protocol. A selected route target is not proof of the provider’s hidden backend build or revision; name it as route evidence unless the provider supplies stronger provenance. The declaration remains a claim until source or provider evidence supports it.
  3. Select cases from practice. Prefer previously completed production or project tasks with independent acceptance evidence. Include materially different cases inside the claimed task population and at least one case likely to expose a characteristic failure. Public benchmarks may supplement this field but cannot replace it. Remove cases so easy that every viable profile passes without exercising the claimed capability, and cases so hard that every run only hits the execution boundary; neither distinguishes the allocation decision.
  4. Name decision-changing evidence before running. For each case state its acceptance conditions, material failure classes, and the observation that would defeat the proposed allocation claim. Avoid criteria that reward verbosity, stylistic resemblance, or test-count inflation. Keep worker-visible acceptance procedural and artifact-oriented. Expected conclusions, seeded defects, reference answers, and semantic comparison criteria remain visible only to the evaluator; otherwise the task measures answer following rather than discovery.
  5. Run matched repetitions. Execute each profile more than once in isolated workspaces. Balance order and avoid parallelism when provider load would add an uncontrolled variable. Retain unsettled runs, retries, latency, usage, selected route identity and any stronger provider provenance, artifacts, verification, and interventions; never discard a failure to make the sample comparable. A run stopped at a declared duration boundary is right-censored evidence: it proves only that the profile did not settle within that envelope. Do not call one boundary hit random instability or estimate its unseen completion time.
  6. Judge blind where judgment is necessary. Hide profile/provider/model identity from the evaluator. Mechanical acceptance may settle deterministic conditions; semantic judgment reports evidence but cannot admit the profile as fact. Keep the judge identity and usage visible outside the blind packet.
  7. Compare variation before averages. Examine within-profile inconsistency, failure modes, and order effects before claiming a between-profile difference. If ordinary variance is as large as the claimed advantage, return inconclusive and change the next probe rather than adding confidence language.
  8. Prepare a Capability Case. Read references/capability-case.md. Link the exact fixture, cases, run records, judge evidence, failures, and resource observations. Treat prompting behavior as evidence about the whole execution profile: retain observed prompt failures and the smallest treatment hypothesis, but do not attribute them to the model alone.
  9. Separate discovery from confirmation. A case used to revise instructions, skills, context, tool descriptions, or completion contracts has become a development case. Hold model, route, harness, tools, and fixture fixed while comparing that prompt treatment; then confirm the revised profile on retained cases that did not teach the treatment. Do not report tuning-case improvement as held-out capability.
  10. Submit deliberately. Only the named human or designated host may accept the candidate claim into a reusable capability profile. An evaluator, runtime, provider, or routing policy cannot approve its own allocation.

Recover only enough state to decide whether a valid evaluation can begin:

Allocation decision the evidence must inform:
Candidate execution profiles and identity revisions:
Real task population and retained acceptance evidence:
Material capability dimensions and failure classes:
Available executor, isolation, usage, latency, and judge evidence:
Existing accepted profile or baseline, if degradation is suspected:
Human or designated host with profile-admission authority:

If there is no allocation decision, real task population, or traceable acceptance evidence, do not manufacture a benchmark. Return the missing evidence and the smallest probe that could obtain it.

  • Do not collapse different task dimensions into one score or leaderboard.
  • Do not infer capability from price, quota, popularity, or provider marketing.
  • Do not let a judge preference erase mechanical failures or divergent runs.
  • Do not silently change prompt, skill, context, tools, permissions, or fallback policy while attributing the result to the model.
  • Do not repeatedly tune on evaluation cases and continue calling them held out.
  • Do not grant provider or spend authority; route configuration remains an environment concern.
  • Route workload estimation and time/money conversion to their owning method after a capability claim exists; evaluation records usage and latency only.
  • Route one proposed code or artifact comparison to its domain review method.

An evaluation is ready for human submission only when the execution identities, task population, fixture and acceptance provenance, repeated raw outcomes, within-profile variance, failure classes, latency and usage, judge identity, alternative explanation, bounded claim, and reopening observation are explicit. If any is absent, retain a probe record rather than a capability profile.

本页有意省略 prompt 的内部装配说明。需要安装、审查或修改这个技能时,请回到 完整源文件