Skip to main content

human-agent-trust-reviewer

Adversarially review the human-approval layer of an agent system for trust exploitation (OWASP Agentic ASI09) — consent fatigue (approval floods that train rubber-stamping; rate/latency as the signals), deceptive, over-polished justifications making a dangerous action look routine, self-reported summaries diverging from the actual action (approvers must see the real diff/blast radius, not the agent's story), dangerous steps bundled in innocuous batches, urgency manipulation, and automation bias as trust accumulates. Counterpart to human-approval-boundary: that skill places the gates; this one attacks their resilience and fixes what folds. Use when reviewing agent approval UX/flows, when approvals feel like a formality, or when an agent's explanations drive human sign-off. Do NOT use to place the gates (human-approval-boundary), audit whether approvals happened (agent-governance-audit), address overreliance on AI content (ai-misinformation-guard), or tier oversight (ai-governance-risk-reviewer).

الانتقال إلى التثبيت

معلومات المصدر

المستودع
ModernNomad-98/Project-Aegis
آخر نشاط في المصدر
١٨ يوليو ٢٠٢٦ في ١٢:٤٦
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.