| name | error-analysis-and-evaluation-discipline |
| kind | skill |
| status | ready |
| description | Review the discipline of the evaluation itself — overt errors classified separately from covert dimensional mismatches, impressionistic criteria rejected unless made operationally checkable, and evidence strength preserved. Use when evaluation criteria, an error typology, or a confidence claim need checking; the source-profile basis they weight against goes to overt-covert-translation-and-equivalence. |
| provenance | {"principles":["P002","P021","P035","P036","P037","P048","P049","P061","P086","P087","P090","P091","P094","P095","P109","P125","P128"],"claims":["C00003","C00004","C00005","C00007","C00011","C00012","C00019","C00021","C00022","C00023","C00024","C00025","C00026","C00030","C00033","C00034"],"evidence":[],"source_anchors":[],"authored_from_digest":"afd640b763a4ac11cf783b42d44fd86aac75cdf298702aa8787dde04fec98b06"} |
Error Analysis And Evaluation Discipline
Purpose
This skill reviews the discipline of the evaluation itself, not the translation it is judging. Overt errors (denotative omissions, additions, substitutions, wrong selections, ungrammaticality) are classified separately from covert (dimensional) mismatches, which are weighted by the source profile and functional component. Global, impressionistic, or contradictory criteria are rejected unless made operationally checkable; analyst judgement is held to argued, evidence-constrained hypotheses; and the strength of the evidence behind every finding is preserved — tentative findings are reported as tentative.
When to use
- Error seriousness is being judged without a specified, text-context-sensitive procedure or across micro-, macro-, and superstructural levels.
- Assessment criteria are global, impressionistic, or contradictory and have not been converted into explicit, checkable textual and contextual tests.
- A tentative, small-sample, or attributed finding is being upgraded into a firm claim, or a hedge is being removed.
- Errors are weighted by a universal hierarchy rather than by their effect on the individual text's ideational and interpersonal functional match.
- Evaluative responsibility for a translation difference is being assigned before obligatory shifts have been separated from optional translator choices.
Procedure
Triage each finding into one of three lenses before writing it up: overt-error classification (denotative/grammatical errors named and kept distinct from judgement calls), covert dimensional mismatch (functional-profile-weighted shifts assessed against the source's profile), and criteria discipline (the evaluation's own criteria, category set, and evidence strength checked before either error type is trusted).
Overt-error classification
- Classify overt errors separately as denotative omissions, additions, substitutions, wrong selections or combinations, ungrammaticality, and questionable acceptability (P090).
- Use micro-, macro-, and superstructural levels to judge error seriousness only as part of a specified, text-context-sensitive procedure (P021).
- Separate obligatory system-driven shifts from optional translator choices before assigning evaluative responsibility for a translation difference (P125).
Covert dimensional mismatch
- State the assumptions and exceptions behind dimensional error judgements, including cultural comparability, intertranslatability, and whether the target has an added special function (P091).
- Weight errors by source-profile priorities, evaluation objective, and functional component, treating denotative mismatches as especially serious when ideational function is central (P094).
- Measure revised-model quality by analogous source and target profile/function analysis while distinguishing dimensional from non-dimensional mismatches (P095).
- Weight errors by their effect on the individual text's ideational and interpersonal functional match rather than by a universal error hierarchy (P128).
- Separate linguistic analysis from social judgement (P048).
- Allow analyst judgement only as argued, evidence-constrained hypotheses, acknowledging that equivalence and TQA retain non-absolute and subjective elements (P061).
Criteria discipline and evidence strength
- Do not use translationese indicators as a direct proxy for translation quality; quality judgments need separate evidence for meaning preservation and pragmatic acceptability (P002).
- Preserve the strength of the evidence (P035).
- Reject global impressionistic or contradictory assessment criteria unless they are converted into explicit, operationally checkable textual and contextual tests (P036).
- Do not accept target-only, reception-only, or response-only evaluations as sufficient when they cannot explain the source-translation relation or distinguish translation from adaptation (P037).
- Account for social, political, ethical, publishing, marketing, reader, and purpose factors without letting them displace linguistic-textual analysis (P049).
- Treat usage norms, reader-response data, market response, and judge samples as insufficient unless their empirical basis and operational procedure are explicit (P086).
- Reject analytic category sets that overlap, lack theoretical grounding, obscure text-context relations, or make target-language evidence inaccessible to evaluators (P087).
- Let a controlled TQA model supply qualitative categories before quantification, especially to reduce stereotype-producing corpus interpretations (P109).
Inputs
- The evaluation's criteria, error classification, weighting rationale, and the strength of the evidence behind each judgement.
- The reasoning offered for the decision under review: the corpus, the orientation, the brief, and any quality claim made.
Output
Per finding: flaw, principle, correction, trade-off, next step — highest-impact first. Reviews the evaluation, not the translation itself or the publication decision.
Anti-patterns to flag
- Treating a translationese or corpus-classifier flag (a stylistic-feature score) as the quality verdict itself, instead of a prompt to gather separate evidence for meaning preservation and pragmatic acceptability (P002).
- Conflating "this finding is tentative" — a small-sample or attributed claim that must stay hedged (P035) — with "all TQA judgement is equally uncertain," as an excuse not to argue a judgement at all; analyst judgement is still an argued, evidence-constrained hypothesis, not a shrug (P061). The two hedges protect different things; neither licenses the other.
- Flattening overt and covert errors onto one severity scale: an overt error (wrong word, ungrammaticality) is classified and counted on its own terms (P090), while a covert dimensional mismatch is only as serious as the centrality of the affected function to the source profile (P094) — a denotative slip where ideational function is peripheral does not automatically outrank a Tenor mismatch where it is central.
- Accepting a criterion because it sounds rigorous — a named category set, a scoring rubric — without checking it is actually operationally checkable (P036), and separately, accepting reader-response or market data as if rigor on the reception side substitutes for explaining the source-translation relation (P037). Sounding rigorous and explaining the source-translation relation are different bars; passing one does not pass the other.
- Correcting for "linguistic analysis doesn't settle every good/bad judgement" (P048) by dropping social, political, or purpose factors from the review entirely, rather than accounting for them without letting them displace the linguistic-textual analysis (P049) — the corrective is additive, not substitutive.
- Assigning responsibility for a translation difference before checking whether it was an obligatory system-driven shift (no reviewable choice) or an optional translator choice (P125), then, once it is optional, ranking it by a generic severity table instead of by its effect on this text's ideational/interpersonal functional match (P128).
References
See ../../references/translation-quality-principles-index.md for the full principle catalogue grouped by skill, and ../../references/translation-quality-evidence-notes.md for how these principles are grounded and kept faithful to the sources.
Provenance
Derived from P002, P021, P035, P036, P037, P048, P049, P061, P086, P087, P090, P091, P094, P095, P109, P125, P128, grounded in the distillation-only source House, Translation Quality Assessment. The frontmatter provenance block lists the exact principle and claim ids, which resolve into principles/principles.yaml and analysis/claims.jsonl.