| name | failure-mode-analysis |
| description | Produces a failure mode and effects analysis (FMEA-lite) with failure
modes, effects, severity ratings, detection methods, and recommended
actions using FMEA methodology. Use when the user asks to perform a
failure mode analysis, identify what could go wrong in a process or
product, create an FMEA, analyze potential failures before they happen,
or prioritize which failures to prevent first.
Do NOT use for general risk assessment (use risk-assessment), process
flow documentation (use process-mapping), or incident post-mortems
(requires incident response process).
|
| license | Apache-2.0 |
| metadata | {"author":"foundry-skills","version":"1.0.0","tags":"analysis planning strategy checklist","category":"business-strategy","subcategory":"operations","depends":"","disclaimer":"none","difficulty":"advanced"} |
Failure Mode Analysis
When to Use
Use this skill when the user presents one of these specific scenarios:
- A process, product, or system is being designed or redesigned and the team wants to identify failure modes before launch -- this is the primary use case for proactive FMEA
- A recurring operational problem is costing money or customer satisfaction and the team wants to understand the full failure landscape, not just the most visible symptom
- A regulated product or process (medical device, automotive component, aerospace system, food production line) requires formal FMEA documentation as part of compliance -- standards such as AIAG-VDA FMEA (4th edition), IEC 60812, and FDA 21 CFR Part 820 mandate this
- A supplier qualification or vendor audit requires an FMEA as supporting documentation for quality assurance
- A post-improvement review needs to validate that an implemented countermeasure actually reduced the RPN as projected (re-FMEA after action)
- A cross-functional team is debating which failure scenarios deserve investment and needs a structured, data-anchored prioritization framework rather than a subjective argument
- A new product introduction (NPI) or change management process requires a Design FMEA (DFMEA) or Process FMEA (PFMEA) as a gate deliverable
Do NOT use this skill when:
- The user wants a broad landscape of risks across categories -- those risks lack a defined function-failure-effect chain and belong in
risk-assessment
- The user needs to map what a process actually does step-by-step before they know what can fail -- complete
process-mapping first, then return here
- Something already failed and the user needs to understand what happened -- that is a retrospective root cause analysis or incident post-mortem, not a proactive FMEA
- The user wants to evaluate a strategic decision (entering a new market, acquiring a company) -- that is scenario planning and strategic risk analysis
- The user needs a quick gut-check on whether an idea is risky -- a full FMEA is a resource-intensive deliverable; for informal risk gut-checks suggest a simple pre-mortem exercise instead
- The process is completely undefined and no subject matter expert is available -- FMEA built without domain knowledge produces dangerously wrong severity and occurrence scores
- The user only wants a compliance checkbox with no intent to act on findings -- this produces a document that harms quality culture without delivering value
Process
Step 1: Define Scope, Boundaries, and FMEA Type
Before generating any failure modes, establish precise scope. An imprecise scope is the most common reason FMEAs become unwieldy or miss critical failures.
- Determine FMEA type. Design FMEA (DFMEA) analyzes how a design could fail to meet its intended function -- used before a product is manufactured. Process FMEA (PFMEA) analyzes how a manufacturing or service process could fail to produce the intended output -- used before a process goes live. System FMEA (SFMEA) analyzes how subsystems interact and fail at the system level -- used for complex multi-component systems. For most business operations questions, PFMEA is the correct type.
- Identify the object of analysis. Name the specific product, component, process, or function. "Accounts Payable Process -- Invoice Receipt through Payment Authorization" is a valid scope. "Finance operations" is not.
- Define the scope boundaries. State explicitly what is included (start trigger, end trigger) and what is excluded. This prevents scope creep and keeps the failure mode list manageable.
- List the process steps or functional elements. For a PFMEA, enumerate every process step in sequence. For a DFMEA, list every design function or component. This list becomes the backbone of the FMEA worksheet.
- Confirm the FMEA team. An FMEA generated by one person is an opinion. A valid FMEA requires cross-functional input: the people who design the step, the people who execute it, the people who are affected by failure, and ideally someone with no stake who asks uncomfortable questions. Document the team roles.
- Collect prior failure history. Ask: What has failed before? What near-misses have occurred? Are there warranty claims, customer complaints, or quality records? Historical failure data is the single best input for realistic occurrence ratings. If none exists, document that explicitly.
- Identify applicable standards, regulations, and customer requirements. Regulatory severity anchors (e.g., FDA, OSHA, ISO 9001, IEC 61508 for safety integrity levels) affect how you calibrate your 1-10 scales and determine mandatory action thresholds.
Step 2: Enumerate Failure Modes Using the Function-Failure-Effect-Cause (FFEC) Chain
This is the analytical core of FMEA. Each entry in the worksheet must complete the full FFEC chain. Incomplete chains produce incomplete -- and therefore misleading -- analyses.
- Function: State what the step or component is supposed to do in specific, measurable terms. Not "process invoices" but "receive, validate, code, and route invoices for payment authorization within 5 business days."
- Failure mode: State how the function fails to be performed. A single function can have multiple failure modes. Common failure mode categories: complete non-performance (function not performed at all), partial performance (function performed incompletely or incorrectly), degraded performance (function performed but below specification), intermittent performance (function performed sometimes but not reliably), and unintended performance (function performs when it should not).
- Failure effect: Describe what happens downstream when this failure mode occurs -- at the immediate step, at the next step, and at the end customer or end state. Most FMEA practitioners miss the downstream cascade. A data entry error at step 1 may only surface as a payment error at step 7.
- Failure cause: Identify the root cause that would produce this failure mode. Use the question "why would this failure mode occur?" not "what is the failure mode?" Common cause categories for process FMEAs: human error (inadequate training, poor interface design, high cognitive load), equipment failure (wear, incorrect calibration, maintenance gap), material or input quality (bad data, defective inputs from upstream), method gap (procedure does not exist or is incorrect), and environmental factors (system outages, resource unavailability).
- Generate at least 3-5 failure modes per major process step. Teams that produce only 1-2 failure modes per step are almost certainly missing partial failures, timing failures, and low-frequency but high-severity failures.
- Explicitly include timing failures. "Done too late," "done too early," and "done in the wrong sequence" are valid failure modes that teams routinely overlook.
- Cap the worksheet at 20-25 rows per analysis session. If the scope requires more, decompose into sub-processes and run separate FMEAs. A worksheet with 50+ rows becomes unactionable.
Step 3: Rate Severity (S) on a Calibrated 1-10 Scale
Severity rates the impact of the failure effect -- not the failure mode itself -- on a 1-10 scale. Severity is a property of the effect, not of how often it happens.
- Calibrate severity to the actual harm, not the emotional reaction. Teams routinely inflate severity because the failure is embarrassing or rate it low because they like the team responsible for that step. Anchor to specific, observable consequences.
- Use this calibrated severity scale:
- 1: No discernible effect. The process continues without interruption or customer awareness.
- 2: Very minor effect. An inconvenience that is corrected before it reaches the next step. No customer impact.
- 3: Minor effect. Internal rework required, minor delay (under 1 hour for a service process). No external impact.
- 4: Low effect. Moderate rework, delay of 1-8 hours, or slight degradation in output quality that the customer may not notice.
- 5: Moderate effect. Customer notices a reduction in quality or a minor delay. Requires customer communication. No SLA breach.
- 6: Moderate-high effect. Deliverable quality noticeably reduced. Customer complaint likely. SLA may be at risk.
- 7: High effect. Customer experiences significant disruption. SLA breach possible. Escalation likely. Substantial internal rework.
- 8: Very high effect. Significant customer-facing failure, potential contract risk, or substantial financial loss (quantify if possible). Recovery is difficult.
- 9: Critical effect -- no safety hazard. System-level failure, regulatory reporting triggered, or customer relationship at serious risk. Near-safety event.
- 10: Critical effect with safety hazard or regulatory violation. Can result in injury, major regulatory action, or systemic customer harm. Immediate action mandatory regardless of RPN.
- For severity 9-10, document the specific safety, regulatory, or systemic harm explicitly. "S=9 because this failure would trigger FDA adverse event reporting under 21 CFR 803" is a valid justification. "S=9 because it feels serious" is not.
- Severity does not change when you add controls. Controls affect detection (D). If the failure effect reaches its destination, the severity is fixed. This is the most common scoring error in FMEA practice.
Step 4: Rate Occurrence (O) on a Data-Anchored 1-10 Scale
Occurrence rates how often this failure mode is expected to occur given current process conditions -- not how often the failure effect is experienced after detection/controls.
- Use this calibrated occurrence scale:
- 1: Failure is essentially impossible. Rate of failure below 0.001% (1 in 100,000 cycles or more). Near-Six Sigma performance.
- 2: Extremely rare. Rate of 0.001-0.01% (1 in 10,000 to 1 in 100,000).
- 3: Very low occurrence. Rate of 0.01-0.1% (1 in 1,000 to 1 in 10,000). Occasional isolated incidents.
- 4: Low occurrence. Rate of 0.1-0.5% (1 in 200 to 1 in 1,000). Occurs once or twice a year in a moderate-volume process.
- 5: Moderate occurrence. Rate of 0.5-2% (1 in 50 to 1 in 200). A familiar problem -- the team knows it happens.
- 6: Moderate-high occurrence. Rate of 2-5% (1 in 20 to 1 in 50). Occurs regularly enough to appear in monthly quality metrics.
- 7: High occurrence. Rate of 5-15% (1 in 7 to 1 in 20). A frequent, visible problem. Team has workarounds.
- 8: Very high occurrence. Rate of 15-30% (1 in 3 to 1 in 7). Failure more common than success in some conditions.
- 9: Near-certain occurrence. Rate of 30-70%. Process is unreliable; failure is the expected outcome in adverse conditions.
- 10: Failure is virtually inevitable with current process. Rate above 70%. This failure mode should trigger immediate process suspension.
- Anchor scores to data whenever possible. Pull from defect logs, quality records, customer complaint data, support ticket history, or process audit findings. Document the data source in the FMEA.
- When no data exists (new process), use expert elicitation with the Delphi method: Collect independent estimates from 3-5 subject matter experts, share the range anonymously, discuss outliers, and converge on a consensus score. Document that the score is an estimate and set a re-assessment trigger after 30-60 operating cycles.
- Do not suppress the occurrence score because of existing controls. Controls belong in detection (D). Rate occurrence as if the current detection method did not exist -- this prevents conflating prevention with detection.
Step 5: Rate Detection (D) on a 1-10 Scale
Detection rates the ability of the current control system to identify the failure mode or its cause before the failure effect reaches the next step or customer. Counterintuitively, a high detection number means poor detection.
- Use this calibrated detection scale:
- 1: Nearly perfect detection. Automated 100% inspection or physical design prevents the failure effect from progressing. Error-proofing (poka-yoke) in place.
- 2: Very high detection certainty. Automated detection catches failures at the point of occurrence >99% of the time.
- 3: High detection. Multiple independent controls catch the failure in >95% of cases before downstream impact.
- 4: Moderately high detection. Systematic checks (checklists, double-check procedures, sampling inspection) catch the failure >85% of the time.
- 5: Moderate detection. Control methods exist but are inconsistent. Failures caught roughly 60-85% of the time.
- 6: Low-moderate detection. Detection relies on individual attention or periodic reviews. 40-60% detection rate.
- 7: Low detection. Controls exist on paper but are rarely applied consistently. Less than 40% of failures are caught.
- 8: Very low detection. Failure is difficult to observe. Catches fewer than 20% of occurrences. Most failures propagate.
- 9: Near-zero detection. No reliable detection method. Failure usually discovered only when the customer reports it.
- 10: No detection. No current control mechanism exists. Failure is invisible until consequence is experienced.
- Classify your current controls into three types and rate accordingly: Prevention controls (prevent the cause from occurring -- these reduce O, not D), detection controls (find the failure before it progresses -- these reduce D), and mitigation controls (limit the effect after failure occurs -- these reduce effective S in some interpretations but do not change D score).
- List controls that actually exist and are actually used. A procedure that exists in a manual but is not followed is not a functioning detection control. Score detection based on real-world practice.
- Error-proofing (poka-yoke) is the gold standard for detection improvement. Physical or digital mechanisms that make it impossible to continue if a failure occurs drive D scores to 1-2. Checklists that depend on human attention rarely achieve better than D=4.
Step 6: Calculate RPN, Apply Priority Rules, and Identify the Critical Path
- RPN formula: RPN = S x O x D. Range is 1 (minimum risk) to 1,000 (maximum risk). The RPN is a relative priority index, not an absolute risk measurement.
- Apply the RPN action threshold. Standard industry thresholds: RPN 200+ requires immediate corrective action with a defined owner and completion date. RPN 100-199 requires a formal action plan within 30 days. RPN 50-99 warrants monitoring and opportunistic improvement. RPN below 50 can be accepted with periodic review. Adjust thresholds to your organization's risk appetite -- regulated industries often set the threshold lower (100+ for immediate action).
- Apply the severity override rule -- the most important rule in FMEA. Any failure mode with S=9 or S=10 requires a mandatory recommended action regardless of RPN. An S=10 failure mode with O=1 and D=1 yields RPN=10 but still demands design change or control implementation because the consequence of occurrence is catastrophic or unsafe.
- Apply the high-S, low-RPN trap check. A failure mode with S=8, O=2, D=2 produces RPN=32 -- which would be "accepted" under a simple threshold rule. But a severity-8 failure happening even rarely warrants a second look. Review all S=7+ items even if RPN is below threshold.
- Rank the worksheet by RPN descending. The top 20% of failure modes by RPN typically account for 80% of risk -- the Pareto principle applies directly here.
- Identify the "critical path" of failures: the sequence of high-RPN failures that, if left unaddressed, would most likely produce the worst end-state outcome. This narrative is more useful to leadership than a table of numbers.
Step 7: Develop Recommended Actions and Project Post-Action RPN
Every recommended action must be specific, dimensioned (which of S, O, or D it targets), owned, and time-bound.
- Prioritize actions that reduce occurrence (O) over actions that improve detection (D). Preventing a failure is always more effective than detecting it after it occurs. Occurrence reduction actions include: redesigning the process step to eliminate the failure cause, mistake-proofing the input or interface, adding a prevention checkpoint before the failure can be generated, and improving upstream supplier or input quality.
- Detection improvement is valid when occurrence reduction is impractical. Adding automated monitoring, implementing 100% inspection at a gate, deploying real-time alerting, or adding a double-check from a second role all improve detection. Detection improvement is faster to implement but does not eliminate the root cause.
- Severity reduction is the rarest action because it requires redesigning the consequence. Sometimes achievable: adding a containment step that limits blast radius (e.g., a rollback mechanism in a software process), reducing the number of customers affected by a single failure, or adding a recovery SOP that limits duration of impact.
- For each recommended action, document:
- The action statement (specific enough that someone could implement it without asking follow-up questions)
- Which dimension it targets (O, D, or S)
- The new expected score for that dimension after implementation
- The projected post-action S, O, D values
- The projected post-action RPN
- The percentage RPN reduction
- The accountable owner (a role, not just a department)
- The target completion date
- Any dependencies on other actions or resources
- Re-assess actions that only reduce D without reducing O as "interim controls." Document them as temporary and schedule an occurrence-reduction follow-up action. A process that only adds inspection without fixing root cause has lower risk but not lower failure rate.
- Assign an action status field: Open, In Progress, Complete, Verified. FMEA is a living document -- each review cycle updates the status and, when actions are complete, re-scores the actual (not projected) post-action S, O, D.
Step 8: Document, Review, and Schedule the Re-Assessment
- Assign a revision number and date to every version. Format: FMEA-[Process Code]-Rev[X.X]-[YYYYMMDD]. This enables traceability when audited.
- Schedule the first re-assessment. For a new process: 30-60 days after go-live, once actual failure data is available to validate or correct occurrence scores. For an existing process: after all high-priority actions are closed, or on a calendar cycle (quarterly for high-criticality processes, annually for stable low-criticality processes).
- Document the "assumed" vs. "confirmed" distinction for all occurrence scores. New processes carry inherent uncertainty. Mark estimated scores with an asterisk or note, and plan to replace them with data-driven scores at first re-assessment.
- Link the FMEA to the control plan. The control plan is the operational document that specifies what checks are performed, by whom, at what frequency. Every detection control in the FMEA should have a corresponding entry in the control plan. This linkage is required for IATF 16949 (automotive) and strongly recommended for ISO 9001 compliance.
- Communicate top findings to leadership in a summary format. Leadership does not need the full worksheet -- they need: the count of S=9/10 items, the top 3 RPNs with owner and due date, and the projected RPN reduction after all actions are closed. One slide or one page is the target.
Output Format
## Failure Mode and Effects Analysis: [Process/Product Name]
**FMEA Type:** [DFMEA / PFMEA / SFMEA]
**Scope:** [Start trigger] through [end trigger] -- excluding [explicit exclusions]
**FMEA Team:** [Role 1], [Role 2], [Role 3] (minimum 3 roles)
**Date Initiated:** [YYYY-MM-DD]
**Revision:** [X.X]
**Next Review Date:** [YYYY-MM-DD]
**RPN Action Threshold:** 100+ (Immediate/30-day action required)
**Severity Override Rule:** All S=9 or S=10 items require action regardless of RPN
**Document ID:** FMEA-[ProcessCode]-Rev[X.X]-[YYYYMMDD]
---
### Executive Summary
**Total Failure Modes Analyzed:** [X]
**Severity Override Items (S=9 or S=10):** [X] failure modes -- immediate action required
**Above RPN Threshold (100+):** [X] failure modes
**Highest RPN Identified:** [X] -- [Failure mode name]
**Lowest RPN Identified:** [X] -- [Failure mode name]
**Projected Total RPN Reduction (after all actions closed):** [X%]
**Critical Path of Risk:** [One paragraph narrative describing the sequence of highest-risk
failures and the worst-case end state if they remain unaddressed. Written for a non-technical
executive reader.]
**Top 5 Priorities:**
| Rank | ID | Failure Mode | RPN | Severity | Action Owner | Due Date |
|------|----|-------------|-----|----------|-------------|---------|
| 1 | FM[X] | [Mode] | [X] | [S/10] | [Role] | [Date] |
| 2 | FM[X] | [Mode] | [X] | [S/10] | [Role] | [Date] |
| 3 | FM[X] | [Mode] | [X] | [S/10] | [Role] | [Date] |
| 4 | FM[X] | [Mode] | [X] | [S/10] | [Role] | [Date] |
| 5 | FM[X] | [Mode] | [X] | [S/10] | [Role] | [Date] |
---
### FMEA Worksheet (Complete)
| ID | Process Step | Function | Failure Mode | Failure Effect (Local / Downstream / End Customer) | Failure Cause | S | O | D | RPN | Score Basis | Current Controls | Control Type | Recommended Action | Action Targets |
|----|-------------|----------|-------------|--------------------------------------------------|--------------|---|---|---|-----|------------|-----------------|-------------|-------------------|----------------|
| FM1 | [Step name] | [Function statement] | [Failure mode] | [Local effect / Downstream effect / Customer effect] | [Root cause] | [1-10] | [1-10] | [1-10] | [SxOxD] | [Data / Expert estimate*] | [Existing controls] | [Prevention / Detection / Mitigation] | [Action] | [Reduce O / Improve D / Reduce S] |
| FM2 | [Step name] | [Function] | [Failure mode] | [Effects] | [Cause] | [S] | [O] | [D] | [RPN] | [Basis] | [Controls] | [Type] | [Action] | [Target] |
*Asterisk indicates estimated score -- to be validated with operational data at first re-assessment.
---
### Detailed Analysis (All Items RPN 100+ and All S=9/S=10 Items)
**[FM-ID]: [Failure Mode Name] -- RPN: [X] | S=[X] O=[X] D=[X]**
| Element | Details |
|---------|---------|
| **Process Step** | [Step number and name] |
| **Function** | [Specific, measurable statement of what this step must accomplish] |
| **Failure Mode** | [Precise description of how the function fails -- be specific enough that a team member could recognize it if they saw it] |
| **Failure Effect -- Local** | [What happens at this step immediately] |
| **Failure Effect -- Downstream** | [What happens at the next affected step(s)] |
| **Failure Effect -- End Customer/Output** | [What the customer or final output experiences] |
| **Failure Cause** | [Root cause -- use "Why?" framing. Not the symptom, the cause.] |
| **Severity (S)** | [X/10] -- [Specific justification tied to the calibrated scale. Reference regulatory or contractual impact if applicable.] |
| **Occurrence (O)** | [X/10] -- [Specific justification with data source, estimated rate, or expert elicitation basis. Mark with asterisk if estimated.] |
| **Detection (D)** | [X/10] -- [Specific justification. Name the control and its real-world reliability.] |
| **Current RPN** | [S x O x D = X] |
| **Current Controls** | [List each control, its type (prevention/detection/mitigation), and its actual effectiveness] |
| **Control Gaps** | [What the current controls fail to address] |
| **Recommended Action** | [Specific, implementable action statement. Who does what, using what method, to achieve what outcome.] |
| **Action Targets** | [Which dimension: Reduce O from X to X / Improve D from X to X / Reduce S from X to X] |
| **Action Type** | [Occurrence reduction / Detection improvement / Severity reduction / Error-proofing] |
| **Owner** | [Specific role -- not a department] |
| **Target Completion Date** | [YYYY-MM-DD] |
| **Dependencies** | [Other actions, resources, or approvals required] |
| **Post-Action S** | [X/10] |
| **Post-Action O** | [X/10] |
| **Post-Action D** | [X/10] |
| **Post-Action RPN** | [X] (reduction of [X%] from current RPN of [X]) |
| **Status** | Open |
---
### RPN Distribution and Action Classification
| RPN Band | Label | Count | Action Required | Review Cadence |
|----------|-------|-------|----------------|----------------|
| 200+ | Critical | [X] | Immediate -- owner and due date this week | Weekly until closed |
| 100-199 | High | [X] | Formal action plan within 30 days | Bi-weekly |
| 50-99 | Medium | [X] | Monitor; improve opportunistically | Monthly |
| 25-49 | Low | [X] | Accept with documented rationale; periodic review | Quarterly |
| Under 25 | Negligible | [X] | Accept; note in FMEA for completeness | Annually |
**Severity Override Items (S=9 or S=10, regardless of RPN):**
| ID | Failure Mode | RPN | S | Status |
|----|-------------|-----|---|--------|
| [FM-ID] | [Mode] | [X] | [9/10] | Action Required |
---
### Action Plan Register
| ID | Failure Mode | Current S/O/D | Current RPN | Recommended Action | Targets | Projected S/O/D | Projected RPN | RPN Reduction | Owner | Due Date | Status |
|----|-------------|--------------|-------------|-------------------|---------|----------------|---------------|---------------|-------|----------|--------|
| FM[X] | [Mode] | [S/O/D] | [X] | [Action] | [O/D/S] | [S/O/D] | [X] | [X%] | [Role] | [Date] | Open |
**Total Current RPN (sum of all items):** [X]
**Total Projected RPN (after all actions):** [X]
**Total Portfolio RPN Reduction:** [X%]
---
### Calibrated Scoring Scales (Reference)
**Severity (S) -- Rates the impact of the failure effect:**
| Score | Label | Business/Service Process Anchor | Product/Manufacturing Anchor |
|-------|-------|--------------------------------|------------------------------|
| 1 | No Effect | No impact; failure corrected before any output is affected | No discernible effect on product function |
| 2 | Very Minor | Internal inconvenience; corrected in < 15 minutes | Very slight cosmetic defect; customer unlikely to notice |
| 3 | Minor | Internal rework; < 1-hour delay; no external impact | Minor cosmetic defect; customer may notice but accepts product |
| 4 | Low | Moderate rework; 1-8 hour delay; quality slightly reduced | Moderate cosmetic defect or very minor functional degradation |
| 5 | Moderate | Customer notices; minor SLA risk; complaint possible | Customer notices functional degradation; product still usable |
| 6 | Moderate-High | Customer complaint likely; SLA at risk; escalation possible | Significant functional degradation; customer dissatisfied |
| 7 | High | SLA breach; significant financial impact; management escalation | Product functions but at substantially reduced level |
| 8 | Very High | Customer relationship at risk; major financial loss; contract risk | Product inoperable; non-safety failure |
| 9 | Critical | Regulatory reporting triggered; near-safety event; systemic harm | Potential safety concern without warning |
| 10 | Catastrophic | Injury, major regulatory action, or data breach with systemic harm | Safety hazard without warning; regulatory violation |
**Occurrence (O) -- Rates frequency of failure mode under current conditions:**
| Score | Label | Approximate Failure Rate | Context |
|-------|-------|------------------------|---------|
| 1 | Essentially Impossible | < 1 in 100,000 | Error-proofed or near-Six Sigma process |
| 2 | Extremely Rare | 1 in 10,000 to 1 in 100,000 | Exceptionally well-controlled process |
| 3 | Very Low | 1 in 1,000 to 1 in 10,000 | Occasional isolated incidents |
| 4 | Low | 1 in 200 to 1 in 1,000 | A few times per year at moderate volume |
| 5 | Moderate | 1 in 50 to 1 in 200 | A familiar, recurring problem |
| 6 | Moderate-High | 1 in 20 to 1 in 50 | Appears in monthly quality metrics |
| 7 | High | 1 in 7 to 1 in 20 | Visible, frequent; team has workarounds |
| 8 | Very High | 1 in 3 to 1 in 7 | Failure common; process reliability poor |
| 9 | Near-Certain | 1 in 1.5 to 1 in 3 | Failure expected in adverse conditions |
| 10 | Inevitable | > 1 in 1.5 | Failure is the normal outcome |
**Detection (D) -- Rates ability of current controls to catch failure before downstream impact:**
| Score | Label | Detection Capability | Typical Control Type |
|-------|-------|--------------------|--------------------|
| 1 | Near-Perfect | Automated 100% inspection; physical poka-yoke prevents progression | Error-proofing device or hard interlock |
| 2 | Very High | Automated detection catches > 99% of occurrences | Automated inspection, sensor-based monitoring |
| 3 | High | Multiple independent controls; > 95% detection rate | Statistical sampling + automated alert |
| 4 | Moderately High | Systematic checks catch > 85% | Formal checklist, double-check procedure |
| 5 | Moderate | Controls exist but inconsistently applied; 60-85% detection | Periodic audit, manual review |
| 6 | Low-Moderate | Relies on individual attention; 40-60% detection | Informal spot-check, supervisory review |
| 7 | Low | Controls rarely applied; < 40% detection | Paper procedure, rarely followed |
| 8 | Very Low | Failure difficult to observe; < 20% detection | No systematic control; ad hoc |
| 9 | Near-Zero | Failure usually discovered only by customer | Customer complaint is the detection mechanism |
| 10 | None | No detection mechanism exists | No control of any kind |
Rules
-
Never omit the full FFEC chain for any failure mode. Every row in the FMEA worksheet must have a Function, a Failure Mode, a Failure Effect (with local, downstream, and end-customer layers), and a Failure Cause. A row with "N/A" or a blank in any of these four fields is incomplete and must be flagged for resolution before the FMEA is used.
-
Severity is fixed by the effect -- it does not change when controls are added. Controls affect detection (D) only. Teams that reduce S when they add a detection control are double-counting the benefit and producing an artificially low RPN. If someone argues that adding a checkpoint reduces severity, redirect them: it reduces detection from 8 to 3, not severity from 8 to 3.
-
Occurrence must be rated as if current detection controls did not exist. Rate how often the failure mode would be generated by the process, not how often it would escape detection. A failure mode that occurs 30% of the time but is caught 90% of the time has O=9 and D=2 -- not O=3. This separation is the fundamental logic of the three-factor model.
-
Any S=9 or S=10 failure mode requires a recommended action, period. This is not optional, even if the RPN is below the threshold. A failure mode that rarely occurs but causes catastrophic harm when it does is not acceptable without a mitigation plan. Document it, own it, close it.
-
Recommended actions must name a specific mechanism, not a category. "Improve training" is not an action. "Create a 20-minute onboarding module covering [specific failure scenario] for all new hires, assessed by a quiz with a minimum 80% pass mark" is an action. Vague actions cannot be implemented, tracked, or verified.
-
Every action must specify which dimension (S, O, or D) it targets and by how much. State the expected post-action score for that dimension and compute the projected RPN. Without this projection, it is impossible to prioritize between two competing actions for the same failure mode.
-
Do not conflate "rare" with "acceptable." Low occurrence scores are not justification for ignoring a failure mode. A failure that occurs once every five years in a safety-critical system may still require mitigation if the severity is 9 or 10. Use the severity override rule.
-
Occurrence scores must cite a basis. Historical defect rate (cite the source), expert elicitation (document the method and participants), or analogous process data (name the process). "The team thinks it happens sometimes" produces a score that cannot be defended in an audit and will not be trusted in a review. Mark all estimated scores with an asterisk and schedule a data-driven re-assessment.
Edge Cases
1. New Process or Product with No Historical Failure Data
When the process or product has never been operated and no analogous process data exists, occurrence scores are entirely speculative. Handle as follows: use structured expert elicitation -- collect independent estimates from at least 3 subject matter experts, present the range to the group without attribution, discuss the highest and lowest scores, and converge on a consensus. Document the method, the participants, and the final rationale in the FMEA. Mark all estimated occurrence scores with an asterisk (*) and add a note: "Estimated -- to be validated after [X] operating cycles." Set an automatic re-assessment trigger at 30 days and 90 days post-launch. In the interim, use conservative (higher) occurrence estimates -- it is safer to over-invest in prevention for a new process than to under-invest because you guessed low.
2. Regulated Industry Requirements (FDA, Automotive IATF 16949, Aerospace AS9100)
Regulatory FMEAs have formal requirements beyond the standard FMEA-lite structure. For FDA Class II/III medical devices, Design FMEA integrates with ISO 14971 risk management -- severity and occurrence thresholds must be linked to acceptable risk criteria in the risk management file. For automotive processes under IATF 16949, use the AIAG-VDA FMEA 4th edition format, which adds a "Structure Analysis" (boundary diagram), a "Function Analysis" (parameter diagram), a "Failure Analysis" (failure net), and splits the traditional FMEA into seven steps with a standardized Action Priority (AP) table replacing raw RPN. The AP table classifies actions as High (H), Medium (M), or Low (L) based on severity-occurrence and severity-detection matrix lookups rather than a single RPN threshold. For aerospace under AS9100, FMEA must be linked to the product's Design History File and controlled as a configuration-managed document. In all regulated contexts, note the applicable standard at the top of the document and confirm which format applies before producing the analysis.
3. Complex Cascading Failures in Highly Automated or Digital Systems
Software processes, cloud infrastructure, and automated workflows fail differently than manual processes. Key differences: failures can cascade across multiple systems in seconds rather than minutes; a single root cause can trigger multiple simultaneous failure modes; the "occurrence" of a software bug may be 100% once the bug is introduced but 0% otherwise -- making standard occurrence scales awkward. Recommended adaptations: identify system dependencies explicitly using a dependency map before enumerating failure modes. Use fault tree analysis (FTA) in conjunction with FMEA -- the fault tree identifies how failure modes combine to produce system-level failures; the FMEA provides the per-component analysis. For digital systems, add "failure mode category" tags: data corruption, service unavailability, performance degradation, security breach, integration failure. Rate occurrence based on deployment frequency and error budget consumption rather than raw failure rate. Detection in digital systems is often strong (monitoring, logging, alerting) -- but validate that alerts are actually responded to, not just generated.
4. Human-Intensive Processes with High Cognitive Load
Processes that rely heavily on human judgment, such as medical triage, financial underwriting, legal review, or customer service escalations, have failure modes dominated by human error. Standard FMEA underestimates human error occurrence because teams are reluctant to assign high occurrence scores to their own team's mistakes. Handle as follows: use a Human Reliability Analysis (HRA) approach to classify human error types -- skill-based errors (automatic actions gone wrong), rule-based errors (wrong rule applied), and knowledge-based errors (inadequate knowledge for novel situation). Rate occurrence using error probability data from the Human Error Probability (HEP) literature for common task types rather than subjective team estimates. For detection, recognize that peer review is D=4 at best, not D=2 -- human checkers miss roughly 15-20% of errors that they were specifically checking for (the theory of signal detection). Design recommended actions toward error-proofing the interface (forcing functions, defaults, confirmations) rather than additional human checking.
5. FMEA for a Process That Has Already Experienced a Major Failure
When the user presents a process that has already had a significant failure and wants to use FMEA to prevent recurrence, the methodology is the same but the starting point is different. The known failure provides a confirmed data point: set the occurrence score for that failure mode at the level that matches the actual failure rate observed. Set severity based on actual impact experienced. Do not let the team minimize these scores because "we've fixed it" -- the FMEA reflects the process as it currently exists, before new actions are implemented. The known failure becomes FM1 or FM2 with a likely high RPN, and the recommended actions for it are the corrective actions already underway. This grounds the FMEA in reality and prevents the team from treating the known failure as less urgent than novel failure modes with uncertain scores.
6. Customer-Facing Processes with External Contractual or SLA Obligations
When failures can trigger contractual penalties, SLA credits, or customer notification obligations, add two columns to the FMEA worksheet: "Contractual Impact" (does this failure trigger an SLA breach, penalty clause, or reporting obligation?) and "Customer Notification Required" (yes/no, and within what timeframe). Failures that trigger contractual impacts should have their severity anchored to the specific contract terms -- a 0.1% SLA breach credit may not warrant S=8, but a breach that triggers contract termination clause absolutely does. For detection, include customer-facing monitoring and SLA tracking systems as controls. Recovery time is relevant to severity: a failure that causes 15 minutes of downtime vs. a failure that causes 4 hours of downtime may have the same failure mode but very different severity scores based on SLA thresholds.
7. Multi-Site or Cross-Functional Processes Where Ownership Is Ambiguous
When a process crosses departmental, team, or geographic boundaries, failure modes that occur at the handoff points are typically the highest-risk and most difficult to own. Handoff failure modes include: information not transferred, information transferred incorrectly, responsibility gap (each side thinks the other owns the step), and timing failure (transfer happens too late). Handle as follows: explicitly add "handoff" as a process step in the FMEA structure. Assign a "responsible interface owner" -- not a department on either side, but a named coordinating role. For occurrence, handoff failures in unstructured processes typically score 6-8 because informal communication is inherently unreliable. Detection is usually 7-9 because handoff failures are invisible until downstream effects appear. Actions should focus on structuring the handoff with a formal protocol, a shared artifact (a handoff document, a system state, a confirmation trigger), and a defined response if the handoff is not received.
Example
Input: "We're launching a new SaaS product for enterprise clients. I want to do a failure mode analysis on our software deployment process. The steps are: environment setup, code deployment, database migration, smoke testing, client-specific configuration, and go-live authorization. We've had a database migration failure in staging that corrupted test data, and we've had configuration errors that affected two beta clients."
Output:
Failure Mode and Effects Analysis: SaaS Enterprise Product Deployment Process
FMEA Type: Process FMEA (PFMEA)
Scope: Software deployment process from environment setup initiation through go-live authorization sign-off -- excludes post-go-live monitoring and incident response procedures
FMEA Team: DevOps Lead, QA Manager, Customer Success Lead, Database Administrator, Product Manager
Date Initiated: [Current date]
Revision: 1.0
Next Review Date: 30 days post-first production deployment
RPN Action Threshold: 100+ (Immediate/30-day action required)
Severity Override Rule: All S=9 or S=10 items require action regardless of RPN
Document ID: FMEA-DEPLOY-Rev1.0-[YYYYMMDD]
Executive Summary
Total Failure Modes Analyzed: 12
Severity Override Items (S=9 or S=10): 2 failure modes -- immediate action required
Above RPN Threshold (100+): 5 failure modes
Highest RPN Identified: 432 -- Database migration corrupts production data
Lowest RPN Identified: 12 -- Environment setup uses wrong server tier
Critical Path of Risk: The highest-risk sequence in this deployment process runs through database migration and client configuration. A migration failure that corrupts production data (RPN 432) can permanently destroy client data and trigger immediate contract review. Even if migration succeeds, an incorrect client-specific configuration (RPN 288) will surface during go-live and create a high-visibility failure at the worst possible moment -- the client's first real use. These two failure modes share a common upstream cause: insufficient pre-deployment validation of environment-specific parameters. A structured pre-deployment checklist with mandatory sign-off by the database administrator and customer success lead, combined with automated configuration validation, would address both failure modes and reduce the combined RPN by approximately 74%.
Top 5 Priorities:
| Rank | ID | Failure Mode | RPN | Severity | Action Owner | Due Date |
|---|
| 1 | FM5 | Database migration corrupts production data | 432 | 9/10 | Database Administrator | 2 weeks |
| 2 | FM8 | Client configuration does not match agreed spec | 288 | 8/10 | Customer Success Lead | 3 weeks |
| 3 | FM9 | Smoke tests pass but don't cover client-specific workflows | 200 | 8/10 | QA Manager | 4 weeks |
| 4 | FM6 | Database migration partially applies -- referential integrity broken | 175 | 7/10 | Database Administrator | 3 weeks |
| 5 | FM11 | Go-live authorization given without CS sign-off | 168 | 7/10 | DevOps Lead | 2 weeks |
FMEA Worksheet (Complete)
| ID | Process Step | Function | Failure Mode | Failure Effect (Local / Downstream / Client) | Failure Cause | S | O | D | RPN | Score Basis | Current Controls | Control Type | Recommended Action | Action Targets |
|---|
| FM1 | Environment Setup | Provision an environment matching production specs | Wrong environment tier provisioned (under-resourced) | Performance issues at smoke test / Deployment aborted or retested / Client go-live delayed | Manual selection from environment catalog without validation against spec sheet | 5 | 3 | 2 | 30 | Expert estimate* | Environment spec checklist | Prevention | Automate environment provisioning from IaC template tied to client contract tier | Reduce O: 3→1 |
| FM2 | Environment Setup | Provision matching production specs | Environment provisioned in wrong region | Data residency violation at go-live / Regulatory breach for EU clients / Contract breach, potential regulatory fine | Region not parameterized in deployment script; manual override possible | 8 | 2 | 3 | 48 | Expert estimate* | Manual region confirmation step | Detection | Hard-code region in client deployment profile; add region validation gate before provisioning proceeds | Reduce O: 2→1, Improve D: 3→1 |
| FM3 | Code Deployment | Deploy correct build artifact to environment | Wrong build artifact deployed (previous version or wrong branch) | Deployment completes but with wrong code / Smoke tests may pass against wrong version / Client receives outdated or incorrect product | No mandatory artifact hash verification; branch selection is manual | 7 | 3 | 4 | 84 | Expert estimate* | Deployment log reviewed post-hoc | Detection | Implement SHA-256 artifact hash verification as a mandatory deployment gate; fail deployment if hash does not match release manifest | Improve D: 4→1 |
| FM4 | Code Deployment | Deploy code without service interruption | Deployment causes unplanned downtime during deployment window | Extended deployment window / Client SLA breach if in agreed window / Client-visible outage |