Use when auditing the cross-cutting "meta" considerations of an AI/ML personnel assessment that apply across every other component — Components 10-12 of the Landers & Behrend (2023) framework: cultural context (power differentials, cross-cultural transfer, community participation), respect (conformance to accepted ethical standards — the Standards, SIOP Principles, OECD Principles, UGAI), and research designs (whether the studies behind every empirical claim are methodologically defensible). Triggers: "cross-cultural AI hiring", "power differentials in algorithm design", "ethical standards conformance audit", "are the studies behind the claims valid", "research design integrity of an AI audit", "community participation in AI design".
Use when scoping or commissioning a psychological audit of an AI/ML personnel assessment — to define which claims the audit will evaluate (validity, utility, lack of bias), establish the auditor's stance and credibility (internal / external / independent), decide formative vs. summative timing and the audience, and settle data/documentation access and disclosure terms. Triggers: "audit an AI hiring tool", "plan an algorithm audit", "bias audit scope", "internal vs external vs independent auditor", "formative vs summative audit", "NYC Local Law 144 bias audit", "what claims should the audit test".
Use when writing up and releasing the results of a psychological audit of an AI/ML personnel assessment — producing a precise, comprehensive technical report for testing professionals AND a layperson-friendly summary for those the predictions affect, establishing the auditor's standards and credibility in the report, and deciding on public release. Triggers: "write the AI audit report", "release the bias audit results", "dual-audience audit report", "should we publish the audit", "auditor credibility statement", "communicate algorithm audit findings".
Use when auditing how an AI/ML personnel assessment is described and how it affects people — Components 7-9 (information & perceptions) of the Landers & Behrend (2023) framework. Covers first-party developer claims (do they honestly and transparently follow from the audit evidence?), second-party effects on those assessed (candidate reactions, justice, false positives vs. false negatives, what is communicated), and third-party understanding (employment-law experts, regulators, community, public). Triggers: "developer marketing claims vs evidence", "candidate reactions to AI hiring", "applicant fairness perceptions", "false positive vs false negative impact", "what do regulators/public think", "transparency of AI hiring claims".
Use FIRST when evaluating, auditing, or debating whether an AI/ML personnel assessment is "fair" or "unbiased" — to define and defend which meaning of fairness/bias applies before drawing conclusions. Covers the three lenses from Landers & Behrend (2023): individual attitudes (distributive/procedural/ interactional justice), legality-ethicality-morality, and technical domain-embedded meanings (statistics vs. machine learning vs. psychometrics). Triggers: "is this AI hiring tool fair/biased", "what does bias mean here", "algorithmic fairness", "disparate impact vs measurement bias in AI", "bias-variance tradeoff", "define fairness for the audit".
Use when auditing the data and foundational design of an AI/ML personnel assessment — Components 1-2 of the Landers & Behrend (2023) framework. Covers input-data population, sampling, range restriction, and incumbent-vs-applicant generalizability; and model design: how the criterion ("ground truth") is defined and its construct validity, why each predictor/feature was included, and whether choices were theory-driven or empirically derived. Triggers: "audit training data", "is the training sample representative", "ground truth validity", "why are these features predictors", "criterion definition in AI hiring", "range restriction in algorithm training data".
Use when auditing how an AI/ML personnel assessment was actually built once its design was set — Components 3-5 of the Landers & Behrend (2023) framework: model development (refinement and cross-validation strategy, documentation), model features (feature engineering — NLP tokens vs. topics, speech-to-text, extracted facial/voice features), and model processes (the estimation algorithm, alternatives explored, and stress tests for bias). Triggers: "audit feature engineering", "k-fold vs holdout vs temporal validation", "NLP bag-of-words bias", "speech-to-text reliability", "how was the model refined", "stress test the algorithm for bias", "model development documentation".
Use when auditing the scores an AI/ML personnel assessment produces — Component 6 of the Landers & Behrend (2023) framework. Covers evaluating the quality of model predictions: reliability (consistency over time and repeated administrations), validity evidence (do scores reflect the claimed constructs and predict the outcome), appropriateness of the cross-validation given generalizability claims, and subgroup differences across protected classes and their intersections. Triggers: "evaluate AI assessment scores", "algorithm reliability and validity", "subgroup differences in algorithm scores", "intersectional bias audit", "does the AI score predict performance", "adverse impact of the model outputs".