Home / Articles / When Models Grade Architecture: Rubric Failures Worth Budgeting For

This article is published in English.

When Models Grade Architecture: Rubric Failures Worth Budgeting For

LLM verifiers invent strict truth, confuse demos with products, and reward boilerplate—design rubrics that respect purpose.

2925 words

Using models to grade architecture answers—and where that breaks

Teams ask LLMs to review systems-architecture responses the way a working architect would. The idea is seductive: scale review, enforce rubrics, catch gaps. In practice, graders invent stricter truth than engineers use, confuse a project’s demo scope with product capability, accept evidenced-but-ridiculous questions, optimize the supplied test rather than the unstated purpose, and leave stale answers after questions change. They also adore plausible boilerplate that sounds senior while saying little.

The system under test

A study setup paired candidate answers about a concrete system with an LLM verifier applying checklists for correctness, evidence, and scope. Human architects disagreed with the model in patterned ways worth cataloging.

Failure 1 — philosophical strictness

Engineers ship under uncertainty; models drift toward absolute definitions of “correct.” A design that is good enough under constraints gets dinged for not proving metaphysical optimality. Rubrics must encode “fit for purpose,” not “true in all possible worlds.”

Failure 2 — project proof versus product capability

If a repository proves a vertical slice, models often claim the whole product already does it. Verifiers must separate evidence about the artifact under test from marketing-scale claims.

Failure 3 — evidenced yet ridiculous questions

A question can cite files and still be the wrong question. Machines optimize against the supplied test; humans notice when the test misses the purpose. Include purpose checks in the rubric.

Failure 4 — fixing questions without re-grading answers

Local edits to prompts without global semantic re-evaluation leave orphan answers that no longer match. Pipeline hygiene: bump a question version and invalidate cached grades.

Failure 5 — boilerplate gravity

Models emit plausible architecture essays—“consider scalability, security, observability”—without engaging the system’s actual bottlenecks. Require citations to concrete components.

250000 Gbps

Design principles for AI verification

  • Rubrics with explicit tolerance for pragmatic engineering.
  • Separate scores for evidence quality versus question fitness.
  • Versioned question banks with forced regrades.
  • Human spot audits on disagreement slices.
  • Ban uncited generic advice in high-stakes grades.

What still works

LLMs help flag missing sections, inconsistent terminology, and absent diagrams when grounded. They are assistants to architects, not replacements for judgment about purpose.

Closing

AI verifying AI fails when the test forgets the human purpose of architecture: decisions under constraints. Build verifiers that respect those constraints, or they will punish the very engineering judgment you hire for.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.

Extended lessons for platform and hiring loops

If you embed LLM graders in interviews or promotion packets, publish the rubric to candidates where ethically allowed, keep humans in the loop on borderline scores, and rotate questions so models cannot memorize a shallow answer key. Track false fail and false pass rates against senior human panels quarterly.

Architecture validation is a socio-technical process. Tools that ignore org constraints, cost envelopes, and operational reality will systematically mis-score practitioners who correctly optimize for those unstated variables. Encode the variables or stop pretending the grader is neutral.

More failure modes worth budgeting for

Failure 6: scoring verbosity as rigor. Failure 7: punishing uncertainty language that experts correctly use. Failure 8: rewarding namedropping frameworks irrelevant to the system. Failure 9: ignoring operational evidence like runbooks and SLOs. Failure 10: drifting rubrics between runs without notice.

Each mode has a mitigation: length caps, credit for calibrated uncertainty, relevance filters, ops evidence requirements, and rubric checksums in grade records.

Practical blueprint

Start with human-written gold answers. Calibrate the model to agree on easy cases. Inspect hard disagreements. Only then automate first-pass grading with mandatory human review above a risk threshold. Never let an irreversible HR decision ride on an un-audited model score.

Depth on purpose alignment

Architecture exists to meet stakeholder goals under constraints. A verifier that cannot name those goals will grade the wrong thing brilliantly. Require goal statements in every question package. Require the answer to restate goals before proposing structure. Grade the coupling between goals and mechanisms, not the eloquence of the mechanisms alone.

Organizational rollout

Pilot on internal design reviews before candidate loops. Measure time saved versus appeal rate. Provide an appeal path that reaches a human architect quickly. Document known model biases. Retrain or replace rubrics when the stack changes—cloud providers, compliance regimes, or traffic shapes.

Recap

LLM verifiers amplify rubric quality. Weak rubrics scaled are scaled weakness. Strong rubrics plus human oversight can accelerate review. The failures above are not reasons to abandon tooling; they are the specification for doing it without fooling yourself.