mit einem Klick
agent-io-skills
agent-io-skills enthält 31 gesammelte Skills von OpenMatter-Network, mit Repository-Berufsabdeckung und Skill-Detailseiten auf SkillsMP.
Skills in diesem Repository
Use when auditing the cross-cutting "meta" considerations of an AI/ML personnel assessment that apply across every other component — Components 10-12 of the Landers & Behrend (2023) framework: cultural context (power differentials, cross-cultural transfer, community participation), respect (conformance to accepted ethical standards — the Standards, SIOP Principles, OECD Principles, UGAI), and research designs (whether the studies behind every empirical claim are methodologically defensible). Triggers: "cross-cultural AI hiring", "power differentials in algorithm design", "ethical standards conformance audit", "are the studies behind the claims valid", "research design integrity of an AI audit", "community participation in AI design".
Use when scoping or commissioning a psychological audit of an AI/ML personnel assessment — to define which claims the audit will evaluate (validity, utility, lack of bias), establish the auditor's stance and credibility (internal / external / independent), decide formative vs. summative timing and the audience, and settle data/documentation access and disclosure terms. Triggers: "audit an AI hiring tool", "plan an algorithm audit", "bias audit scope", "internal vs external vs independent auditor", "formative vs summative audit", "NYC Local Law 144 bias audit", "what claims should the audit test".
Use when writing up and releasing the results of a psychological audit of an AI/ML personnel assessment — producing a precise, comprehensive technical report for testing professionals AND a layperson-friendly summary for those the predictions affect, establishing the auditor's standards and credibility in the report, and deciding on public release. Triggers: "write the AI audit report", "release the bias audit results", "dual-audience audit report", "should we publish the audit", "auditor credibility statement", "communicate algorithm audit findings".
Use when auditing how an AI/ML personnel assessment is described and how it affects people — Components 7-9 (information & perceptions) of the Landers & Behrend (2023) framework. Covers first-party developer claims (do they honestly and transparently follow from the audit evidence?), second-party effects on those assessed (candidate reactions, justice, false positives vs. false negatives, what is communicated), and third-party understanding (employment-law experts, regulators, community, public). Triggers: "developer marketing claims vs evidence", "candidate reactions to AI hiring", "applicant fairness perceptions", "false positive vs false negative impact", "what do regulators/public think", "transparency of AI hiring claims".
Use FIRST when evaluating, auditing, or debating whether an AI/ML personnel assessment is "fair" or "unbiased" — to define and defend which meaning of fairness/bias applies before drawing conclusions. Covers the three lenses from Landers & Behrend (2023): individual attitudes (distributive/procedural/ interactional justice), legality-ethicality-morality, and technical domain-embedded meanings (statistics vs. machine learning vs. psychometrics). Triggers: "is this AI hiring tool fair/biased", "what does bias mean here", "algorithmic fairness", "disparate impact vs measurement bias in AI", "bias-variance tradeoff", "define fairness for the audit".
Use when auditing the data and foundational design of an AI/ML personnel assessment — Components 1-2 of the Landers & Behrend (2023) framework. Covers input-data population, sampling, range restriction, and incumbent-vs-applicant generalizability; and model design: how the criterion ("ground truth") is defined and its construct validity, why each predictor/feature was included, and whether choices were theory-driven or empirically derived. Triggers: "audit training data", "is the training sample representative", "ground truth validity", "why are these features predictors", "criterion definition in AI hiring", "range restriction in algorithm training data".
Use when auditing how an AI/ML personnel assessment was actually built once its design was set — Components 3-5 of the Landers & Behrend (2023) framework: model development (refinement and cross-validation strategy, documentation), model features (feature engineering — NLP tokens vs. topics, speech-to-text, extracted facial/voice features), and model processes (the estimation algorithm, alternatives explored, and stress tests for bias). Triggers: "audit feature engineering", "k-fold vs holdout vs temporal validation", "NLP bag-of-words bias", "speech-to-text reliability", "how was the model refined", "stress test the algorithm for bias", "model development documentation".
Use when auditing the scores an AI/ML personnel assessment produces — Component 6 of the Landers & Behrend (2023) framework. Covers evaluating the quality of model predictions: reliability (consistency over time and repeated administrations), validity evidence (do scores reflect the claimed constructs and predict the outcome), appropriateness of the cross-validation given generalizability claims, and subgroup differences across protected classes and their intersections. Triggers: "evaluate AI assessment scores", "algorithm reliability and validity", "subgroup differences in algorithm scores", "intersectional bias audit", "does the AI score predict performance", "adverse impact of the model outputs".
Use when considering how candidates react to an AI/ML selection tool and what is communicated to candidates and stakeholders about it — Concerns 9-10 of Tippins, Oswald & McPhail (2021). Covers applicant reactions and their tenuous link to behavior, the faking-vs-training question for video interviews, pitfalls in reaction metrics, and what information can/should be shared with unsuccessful applicants and other stakeholders. Triggers: "candidate reactions to AI hiring", "applicant perceptions video interview", "faking vs training interview", "what to tell rejected candidates", "explain AI hiring decision", "what to share with stakeholders about selection".
Use when an AI/ML selection tool uses data the candidate does not control or did not knowingly provide — scraped social-media/Internet data, or incidental data like facial micro-expressions, voice, and appearance — Concern 8 of Tippins, Oswald & McPhail (2021). Covers the loss of applicant control, job-irrelevance and "is it fair," reputation-scrubbing services and adverse impact, the absence of a clear legal/ethical rule, informed consent (Illinois AIVI Act), and the range of policy approaches. Triggers: "scraped social media hiring", "data outside applicant control", "facial appearance in hiring", "is it fair to use this data", "informed consent for AI hiring data", "online reputation scrubbing".
Use when an AI/ML selection tool uses "dynamic" models or norms that update frequently (sometimes after every administration), or when deciding how often to revalidate and update norms — Concern 7 of Tippins, Oswald & McPhail (2021). Covers the real-change-vs-instability dilemma, technical-report/documentation updates, score adjustments and grandparenting, disparate-treatment risk from candidates evaluated on different variables, applicant-pool shifts affecting validity/range restriction, and revalidation cadence. Triggers: "dynamic models hiring", "algorithm updates after every administration", "how often revalidate AI", "norms updating", "grandparenting test scores", "candidates scored on different variables", "continuous validation".
Use when framing the profession-level response to AI selection tools, or orienting a project to the governing standards — the "Call to Action" of Tippins, Oswald & McPhail (2021). Covers the Principles and Standards as the two guiding documents, the argument that SIOP should develop interpretive guidance APPLYING the Principles to technologically enhanced assessments (not rewrite them), the need for interdisciplinary collaboration, and the warning against letting practice reach "escape velocity" from scientific, legal, and ethical moorings. Triggers: "how should the profession respond to AI hiring", "extend the Principles to AI", "interpretive guidance for AI assessments", "what standards govern AI selection", "I-O psychologists role in AI hiring", "call to action AI selection".
Use when checking whether an AI/ML selection tool is grounded in an adequate analysis of work and is demonstrably job-related — Concerns 2-3 of Tippins, Oswald & McPhail (2021). Covers whether a job analysis is necessary (even with a strong criterion-related relationship), acceptable forms and rigor, O*NET and competency-model limits, collecting task importance ratings, SME-judgment agreement, and the legal meaning of job relatedness. Triggers: "does the AI tool need a job analysis", "is this algorithm job-related", "competency model vs job analysis for AI", "O*NET as job analysis", "job relevancy of scraped predictors", "Guardians job analysis".
Use when evaluating whether the machine-learning methodology behind a selection tool is appropriate and interpretable — Concern 4 of Tippins, Oswald & McPhail (2021). Covers ML interpretability and the "black box," explainable AI (XAI), evaluation metrics (MSE, confusion matrix, ROC/AUC), the high variable-to-case ratio in big data, the difficulty of comparing ML results to traditional methods, and the I-O psychology education gap. Triggers: "is the ML methodology appropriate", "black box hiring model", "explainable AI selection", "ROC AUC confusion matrix", "how to evaluate a machine learning model", "compare ML to regression validity", "I-O psychologists machine learning training".
Use when an AI/ML selection tool uses predictors with no clear theoretical or job-analytic rationale — scraped data (resumes, social media, emails, the Internet), voice/facial features, or opaque big-data correlations. Covers the debate over whether predictors need a theoretical basis, proxy-variable risk (e.g., ZIP code for race), and how the presence or absence of adverse impact changes the analysis. Maps to Concern 1 of Tippins, Oswald & McPhail (2021). Triggers: "atheoretical predictors", "scraped data hiring", "why does this variable predict", "proxy variables in AI hiring", "is a correlation enough", "predictor with no rationale".
Use when evaluating the reliability (score consistency/stability) of an AI/ML selection tool — Concern 6 of Tippins, Oswald & McPhail (2021). Covers reliability as an absolute requirement, what stability means for AI scores, evidence that machine scoring can be as or more reliable than human scoring, the questionable reliability of facial-emotion analysis (including across skin tone, disability, and altered features), and confounds from individual differences in the data generated (e.g., extraversion/verbosity). Triggers: "reliability of an AI assessment", "are the scores stable", "facial emotion recognition reliability", "machine-scored interview reliability", "test-retest for AI hiring", "verbosity confound".
Use when evaluating the professional-ethics obligations around an AI/ML personnel selection tool under the APA Ethics Code — Concern 11 of Tippins, Oswald & McPhail (2021). Covers Ethics Code Section 9 (9.01 Bases for Assessments, 9.02 Use of Assessments, 9.03 Informed Consent), the difference in consent standards for researchers (8.05) vs. those employing tools, and how reliability, validity, and fairness are intertwined with ethical duties. Triggers: "APA ethics AI hiring", "is it ethical to use this assessment", "informed consent assessment", "ethics code section 9", "psychologist responsibility AI selection", "implied consent job applicant".
Use when assessing the legal and regulatory exposure of an AI-based / technologically enhanced personnel selection tool (primarily U.S., with global notes). Covers the Uniform Guidelines on Employee Selection Procedures, Title VII disparate impact and the job-relatedness/business-necessity defense, the OFCCP's 2019 position on AI, the Guardians content-validation case, the Illinois AI Video Interview Act and other state laws, and why "no adverse impact" does not equal "valid." Triggers: "is this AI hiring tool legal", "Uniform Guidelines and AI", "disparate impact algorithm", "OFCCP AI", "Illinois AI Video Interview Act / BIPA", "do we need to validate the algorithm", "adverse impact vs validity".
Use FIRST when evaluating, classifying, or comparing any AI-based or technologically enhanced personnel selection tool — to separate the three independent things it combines: technologies, data, and algorithms (Tippins, Oswald & McPhail, 2021). Establishes that a technology is never "universally valid," that data range from intentional to incidental, and that ML effectiveness depends more on data quality than algorithm choice. Triggers: "evaluate an AI hiring tool", "is this video-interview/game/social-media tool valid", "AI vs ML vs deep learning", "supervised vs unsupervised selection", "big data hiring", "what does this technology actually measure".
Use when determining what validity evidence an AI/ML selection tool needs and whether it has it — Concern 5 of Tippins, Oswald & McPhail (2021). Covers validity as the legal/business sine qua non when adverse impact exists, AI tools as "tests" requiring validity, criterion-related vs. content strategies and why content validation is hard for AI, overall model fit (R-squared) as sometimes the only basis, comparative data for less-adverse alternatives, and the minimum documentation requirements. Triggers: "what validity evidence does the AI tool need", "validate an algorithm", "content validation for AI", "R-squared as validity", "document AI validation", "is overall model fit enough".
Use when preparing the administration documentation / manual for an operational selection procedure — the materials administrators and users need to administer, score, interpret, secure, and communicate the procedure consistently. Covers administrator qualifications, the testing environment, scoring/interpretation, security, candidate communications and feedback, reassessment, nonstandard administrations, data retention, and review/updating. Triggers: "administration manual", "test administration documentation", "test security", "candidate feedback", "retest policy", "reassessment", "data retention for test scores", "proctoring / unproctored internet testing".
Use when assessing candidates with disabilities or from different linguistic/ cultural backgrounds — deciding and documenting selection-procedure accommodations vs. modifications, preserving score comparability and construct measurement, and handling translation/adaptation. Covers the accommodation-vs- modification distinction, the candidate dialog, score handling, documentation, legal limits, and consistency with operational use. Triggers: "test accommodation", "disability accommodation", "modify a test", "extended time / Braille / screen reader", "construct-irrelevant barrier", "translate a test", "linguistic/cultural background", "ADA testing".
Use when building validity evidence by linking selection-procedure content to the work domain — demonstrating that the procedure samples important work behaviors, activities, and/or worker KSAOs defined by an analysis of work. Covers defining the content domain, SME qualifications and linkage judgments, sampling the domain, fidelity/specificity, competency-model foundations, and evaluating content evidence. Triggers: "content validity", "content-based strategy", "work sample", "job knowledge test", "link test content to the job", "SME linkage ratings", "no criterion data available".
Use when designing, conducting, or evaluating a criterion-related validity study — demonstrating an empirical relationship between selection-procedure (predictor) scores and work-relevant criteria. Covers predictive vs. concurrent designs, criterion development (relevance/contamination/deficiency/ reliability/bias), predictor choice, participant sampling, statistical power, data analysis, corrections for range restriction and unreliability, and combining predictors/criteria. Triggers: "criterion validity", "predictive/ concurrent study", "validity coefficient", "correct for range restriction", "criterion measure", "is the test related to performance".
Use when evaluating fairness and bias of a selection procedure — distinguishing the several meanings of "fairness," testing for predictive bias (differential prediction via moderated regression), and examining measurement bias (DIF, item sensitivity review). Covers what subgroup-mean differences do and don't imply, when bias analyses are warranted, and the statistical pitfalls. Triggers: "adverse impact vs bias", "differential prediction", "predictive bias", "measurement bias", "DIF analysis", "is the test fair / biased", "subgroup differences", "item sensitivity review".
Use when justifying a selection procedure with validity evidence gathered elsewhere instead of (or alongside) a local study — via transportability, synthetic/job-component validity, or meta-analytic validity generalization (VG). Covers when each applies, the work-analysis link required, moderators, and limits on generalization. Triggers: "validity generalization", "use existing/meta-analytic evidence", "transport a study", "synthetic validity", "job component validity", "do we need a local study", "borrow validity".
Use when examining the internal structure of a selection procedure as supporting validity evidence — the relationships among items/components and how they conform to the intended conceptual framework (unidimensional vs. multidimensional). Covers dimensionality, factor-analytic models, when coefficient alpha is and isn't appropriate, and validity of subscores. Triggers: "internal structure", "dimensionality", "factor analysis / CFA", "coefficient alpha", "is this scale unidimensional", "subscale validity".
Use when deciding how to combine selection procedures and turn scores into decisions — compensatory vs. multiple-hurdle models, cutoff scores, banding, rank-order/top-down selection, norms, and communicating effectiveness via expectancy charts and utility. Covers the validity/diversity tradeoffs and the documentation each choice requires. Triggers: "cutoff score", "banding", "rank order vs cutoff", "compensatory vs multiple hurdle", "combine test scores", "set a passing score", "utility analysis", "expectancy chart", "norms".
Use when writing or reviewing the technical validation report that documents a selection-procedure validation effort — the record that lets another competent professional understand, evaluate, replicate, and draw independent conclusions about the work. Covers every required section and a reusable outline. Triggers: "technical report", "validation report", "document the validation study", "write up the validity study", "what goes in the validation report".
Use when planning or scoping a personnel-selection validation effort BEFORE data collection — to define the organization's needs, objectives, and constraints, specify the proposed uses and inferences, decide whether existing evidence suffices or new evidence is needed, and choose a validation strategy (criterion-related, content, internal structure, or generalization). Triggers: "plan a validation study", "which validation strategy", "do we need a local validity study", "scope a selection project", "justify a test for this job".
Use when analyzing work to support a selection effort — to identify worker requirements (KSAOs/competencies), define the work content domain, or develop job-relevant criteria. Covers choosing methods, level of detail, SME sampling, competency modeling, and documenting results. Triggers: "job analysis", "work analysis", "define KSAOs/competencies", "competency model", "what should the test measure", "identify the criteria for this job", "SME workshop".