FMEA, Explained So Your Whole Team Gets It.
Failure Mode and Effects Analysis is the closest thing quality has to a time machine — it lets you fix failures before they happen. Whether you're a student meeting FMEA for the first time or a quality head migrating to AIAG-VDA, this guide covers the failure chain, the 7 steps, DFMEA vs PFMEA, and Action Priority — in plain language, with the tips we teach in real workshops.
What Is FMEA, Really?
FMEA is a structured team exercise built on three questions: What could go wrong? What would happen if it did? Why would it happen? Each answer becomes a failure chain, each chain gets rated for Severity, Occurrence and Detection, and the team spends its energy where the risk genuinely is. Born on 1940s military programs, matured through aerospace and automotive, FMEA is now one of the five core tools of IATF 16949 — and since 2019, the AIAG-VDA FMEA Handbook is the harmonized global reference.
Master One Idea First: The Failure Chain
Ninety percent of FMEA confusion is level confusion — arguing whether something is a mode, a cause or an effect. The answer is always: it depends where you're standing. Pick the focus element, and the chain sorts itself out.
The 7 Steps, One by One
The AIAG-VDA handbook organizes FMEA into seven steps — three to understand the system, three to find and reduce risk, one to communicate. Here's each step with the tip that saves teams the most pain.
Planning and Preparation
Define what this FMEA covers before anyone opens the form. The handbook gives you the 5T memory aid: InTent (why are we doing it), Timing (when — early enough to matter), Team (who), Task (which analysis, which scope), Tools (which software/format). Decide what's in scope and — just as important — what isn't. Pull the lessons learned, warranty data and previous FMEAs for similar products.
The single best predictor of FMEA quality is whether it started before or after design freeze. An FMEA scheduled after the tooling order is documentation, whatever it says on the cover.
Structure Analysis
Break the thing you're analyzing into its parts. For a DFMEA: system → subsystems → components. For a PFMEA: the process → process steps (stations) → work elements, using the 4M (man, machine, material, method) to find them. The output is a structure tree — and each element in it will become a "focus element" the later steps analyze.
Draw the tree on a whiteboard before touching software. When teams start in the form, they skip structure entirely and end up analyzing whatever the first column suggests.
Function Analysis
Give every element in the structure its functions — what it must do, with measurable requirements. "Bolt joins bracket to frame with 25±3 Nm clamp load" is a function; "bolt is good quality" is not. Functions link across levels: the component's function enables the subsystem's, which enables the system's. This chain is what makes the failure analysis rigorous instead of brainstormed.
Write functions as verb + noun + measurable requirement. If you can't measure it, you can't rate its failure — and step 4 will be opinions.
Failure Analysis
Now negate the functions: for each focus element, how could it fail to deliver its function (failure mode)? What happens upstream to the customer (failure effect)? What in the lower level causes it (failure cause)? This creates the failure chain — Effect ← Mode ← Cause — and failure networks that connect across the structure levels.
The most common confusion in every workshop: mode vs cause vs effect depends on which level you're standing on. One level's cause is the level below's mode. Pick the focus element first and the confusion disappears.
Risk Analysis
Rate each chain: Severity (of the effect), Occurrence (of the cause, considering your prevention controls) and Detection (of your detection controls), each on defined 1–10 tables. Then read the Action Priority table: High, Medium or Low priority for action. Document the current prevention and detection controls honestly — the ratings are only as real as the controls behind them.
Rate Occurrence with the prevention control in front of you, not from memory. "We have a design rule for that" — show it. If it can't be shown, it doesn't reduce Occurrence.
Optimization
The step that justifies all the others: define actions that actually reduce risk — prevention actions to cut Occurrence, detection actions to improve Detection, design changes to cut Severity where possible. Every action gets an owner and a date, and after implementation the chain is re-rated with evidence. High AP items need action; Medium needs action or a documented justification.
Track FMEA actions in the same system as your other actions — a separate FMEA action list is where good intentions go to expire. This is also where software beats spreadsheets outright.
Results Documentation
Tell the story. Summarize scope, method, the high-risk items and what was done about them — for internal management and, in appropriate depth, for the customer. The FMEA itself remains a living document: field failures, 8Ds and process changes reopen it.
Keep a one-page FMEA summary per program: top risks, actions closed, residual Highs. When the customer auditor asks "show me how you manage design risk," that page answers in thirty seconds.
DFMEA vs PFMEA
Same method, different object: the DFMEA interrogates the design, the PFMEA interrogates the process — and the special characteristics discovered in the first flow into the second. Full comparison in our DFMEA vs PFMEA guide.
Action Priority: Why RPN Had to Go
For decades teams multiplied S×O×D into an RPN and argued about thresholds. The AIAG-VDA handbook replaced that arithmetic with a lookup logic that weights Severity first and answers the only question that matters: how urgently does this chain need action?
FMEA-MSR: The Third Analysis
Electronics changed the failure landscape: a sensor can fail in the customer's hands, and what matters then is whether the system notices and responds safely. The handbook's Supplemental FMEA for Monitoring and System Response (FMEA-MSR) analyzes exactly that — the failure during operation, the diagnostic monitoring that detects it, and the system response that keeps the user safe.
If you build ECUs, sensors, or anything with onboard diagnostics for automotive customers, expect FMEA-MSR to appear in your customer requirements alongside DFMEA.
Where FMEAs Go to Live (or Die)
An FMEA earns its keep only if it stays alive: actions tracked to closure, ratings re-scored with evidence, special characteristics flowing into the control plan and checksheets, and field failures reopening the analysis. That lifecycle is exactly what dies in spreadsheets — and exactly what FAST FMEA software automates: AIAG-VDA structure, AP calculation, action tracking with escalations, and revision control that auditors can follow.
Honest scoping note: software accelerates a team that understands the method. It cannot replace the thinking — which is why this page exists.
FMEA: Frequently Asked Questions
Run FMEAs That Change Designs, Not Just Fill Forms
Get the complete AIAG-VDA FMEA checklist, or see structure trees, Action Priority and linked control plans running live in FAST FMEA software.