Fail-Operational EV Power Electronics: When Fail-Safe Shutdown Is Not Enough
Quick Overview
Problem: A drivetrain that passes every fault test on the bench can lose controllability in the vehicle, because de-energising the power stage is not the safe state for every fault — traction power electronics directly determine available propulsion and regenerative-braking torque.
Common failure points: A rapid drop in drive torque when an inverter trips, an abrupt regenerative-to-friction brake handoff, low-voltage undervoltage or resets when a DC/DC fault meets weak buffering, non-redundant control channels with no independent comparison, and degraded modes that are specified but under-validated.
Where it appears: Traction inverters, DC/DC converters, integrated drive units, 400 V and 800 V drivetrains, and platforms where power electronics are coupled to steering, braking, or higher automation levels.
Engineering focus: Detecting and reacting to a fault inside its fault-tolerant time interval, deciding per hazard whether the safe transition is controlled degradation or shutdown, redundancy where loss of function is unrecoverable, and ISO 26262 validation of every degraded state.
An Illustrative Integration Scenario
Consider an EV drivetrain that passes bench fault tests but behaves differently during vehicle-level integration. On the bench, inverter efficiency is in spec, the DC/DC holds its rail, and every injected fault trips protection exactly as designed.
Put the same unit in a vehicle and reactions that looked correct in isolation can read differently in motion. A gate-driver fault that trips the inverter reduces available torque quickly — clean on a dry, straight road, less benign mid-corner or on a low-grip surface, where an uncoordinated deceleration becomes a disturbance the stability system has to absorb. And if a DC/DC fault meets a low-voltage architecture without enough buffering or isolation, the controllers that coordinate steering and braking can see undervoltage at the worst possible moment.
The protection mechanisms may have worked exactly as designed. What is missing is a system-level response after the initial fault — an answer to a different question: what should the system do while the vehicle is still moving and something has already failed?
That is the gap between a fail-safe drivetrain and a fail-operational one.
The Assumption That Breaks
Teams typically assume: de-energising the power stage is the safe state for every credible fault.
In reality: fail-safe design aims to transition the system to a state that does not create unreasonable risk, and that state is defined by hazard analysis, not by what protects the silicon. For some EV faults, de-energising is exactly that safe state. For others — where traction power electronics directly determine available propulsion and regenerative-braking torque — removing torque mid-corner, cutting regeneration mid-stop, or starving the low-voltage domain can create more risk than the original fault. The design question shifts from “how do we reach a safe state?” to “how do we stay controllable long enough to reach one?”
Why Fail-Safe Breaks Down
An immediate shutdown is not always the safest transition. When a traction inverter stops switching, available drive torque can fall rapidly after the gate signals are disabled. On a low-grip surface that uncoordinated change can provoke the instability the safety system is meant to prevent, and it can sharply reduce regenerative braking, requiring coordinated compensation by the friction-brake system. Whether shutdown is the correct reaction depends on the hazard, which is what a functional-safety analysis under ISO 26262 is for.
The low-voltage domain can propagate a single fault. Most EVs step the high-voltage battery down through a DC/DC converter to one or more low-voltage buses, commonly 12 V and, in some architectures, 48 V, powering control units, buses, and actuators. A DC/DC fault does not automatically collapse those buses — vehicles usually carry an LV battery or other buffering. But if the low-voltage architecture lacks sufficient buffering or fault isolation, a DC/DC failure can propagate into undervoltage or reset conditions in the controllers that steering and braking depend on.
Detection, takeover, and continuation are three different problems. A single non-redundant control channel cannot provide an independent comparison of its own output. Adding a lockstep core does not by itself make a system fail-operational: lockstep improves diagnostic coverage, and a lockstep mismatch commonly drives the MCU into a safe state rather than handing control to a second core. Continued operation needs an independent controller or safety channel that can take over or command a controlled degradation, and, for functions that must keep running, a genuinely redundant control path.
Degraded modes are often under-validated and under-coordinated. A recurring integration risk is that degraded states receive less dynamic validation than normal operation and shutdown, so their transitions overshoot or trip a second protection when finally used. And because fail-operational behavior is not local, the inverter, brake controller, and vehicle control unit must share a consistent view of available torque; without that, a well-behaved local fallback still produces a badly behaved vehicle.
The System Underneath
HV battery → HV DC bus → DC-link capacitor → traction inverter (gate drivers → power switches) → motor → torque → vehicle dynamics — with a branch: HV bus → DC/DC converter → 12 V / 48 V buses → steering, braking, VCU, communications.
The power stage is not one stage in that path; it is the medium the whole drivetrain acts through, and it feeds a second path that keeps the control domain powered.
A single-leg short in the inverter is caught first by local hardware protection: desaturation monitoring in the gate driver detects certain short-circuit conditions and initiates a protective turn-off within microseconds — on the order of a few microseconds for some SiC devices, shorter than typical silicon IGBTs. Detection is not the same as recovery: a device that fails stuck-short may still require separate isolation, and that fast local reaction is not a system-level decision. Whether the drivetrain can then continue in a reduced mode is fixed at design time — only if the power-stage and motor topology were built for phase or leg isolation and fault-tolerant operation. If they were not, the correct outcome is the trip.
Where continuation is designed in, coordinating it across the inverter, brake, and vehicle-control units is the same system-level ownership problem that the move to centralized and zonal compute architectures forces onto every safety-relevant function: the fast local reaction is necessary, but a safe vehicle-level outcome only exists if the slower, system-level context was designed rather than assumed.
Failure Patterns
These are illustrative patterns, not accounts of a specific project.
Pattern 1 — torque step on trip. A single-phase fault trips the inverter at speed. Fail-safe logic stops all switching and drive torque falls quickly, uncoordinated with the brakes. On a low-grip surface the stability controller must recover a disturbance that a torque-limited degraded mode, if one existed and were validated, would have avoided.
Pattern 2 — low-voltage propagation. A DC/DC fault meets an LV architecture with thin buffering or no fault isolation. Dependent controllers see undervoltage; some reset. The trigger is not the converter fault alone but the absence of buffering and isolation between it and the safety-relevant loads.
Pattern 3 — missing sensor fallback. The rotor-position signal from the resolver is lost. Where the motor, operating point, available measurements, and validated safety concept permit an observer-based estimate, the drive could hold reduced torque through a controlled transition; where they do not — low speed is a common limit — the trip is the honest answer. The failure is treating the fallback as universal.
EV Power Electronics and Functional-Safety Engineering
Fail-operational behavior in an EV drivetrain — a validated degraded trajectory instead of a bare trip, redundancy only where loss is unrecoverable, and coordinated fault handling across the inverter, brake, and vehicle-control units — is an architecture-and-validation problem, not a protection-circuit bug. Promwad develops power electronics and control firmware for traction inverters, DC/DC converters, motor drives, and battery systems, with functional-safety work up to ASIL-C, HIL testing, and Automotive SPICE-traceable development.
Engineering Experience Across Power and Motor-Control Platforms
Case Evidence: What Promwad's Public Work Supports
Rather than a single narrative, here is what the public record actually shows, and where the fail-operational step sits relative to it.
Power-stage and inverter architecture. A 48 V mild-hybrid inverter architecture project covered architecture selection, PLECS/Simulink simulation, thermal calculation, and verification against driving profiles — the discipline a fault-tolerant power stage is built on, though this project scope did not include fail-operational degraded modes. See Power Inverter Architecture Design.
Independent monitoring and cross-checking. A dual-MCU battery-management unit with an independent safety MCU and cross-monitoring was built to railway SIL 2. It is adjacent proof of the detection-and-monitoring pattern fail-operational control relies on — not, on its own, evidence of an automotive fail-operational traction inverter. See Dual-MCU Railway BMU Architecture.
Sensor integrity under functional safety. An automotive current sensor developed to ASIL-C demonstrates functional-safety design and sensor integrity in the power path — a building block for fault handling, distinct from a full degraded-operation concept. See ASIL-C Current Sensor Design.
Motor control. FOC with sensored and sensorless control for PMSM, BLDC, and ACIM is documented in Motor Control Engineering — the control layer where reduced-phase and observer-based fallbacks would be implemented.
Taken together these are the component competencies a fail-operational EV drivetrain requires. The engineering step this article describes — combining them into validated degraded modes with vehicle-level fault-injection — is the work, not a claim that a finished fail-operational drivetrain already ships.
Working Approach
Step 1: Define the safe transition per hazard, not per component. For each credible fault, use hazard analysis to decide whether the safe transition is controlled degradation or shutdown, then budget detection plus reaction time to fit inside the fault-tolerant time interval before the fault becomes hazardous. That budget is the design input that power-electronics architecture work should carry from the start, not a commissioning afterthought.
Step 2: Add redundancy only where loss of function is unrecoverable. Multi-phase or segmented power stages can continue with a faulted phase if — and only if — the topology was designed for isolation; a buffered or segmented low-voltage path keeps safety-relevant controllers powered; an independent monitor or second controller provides the takeover that lockstep detection alone does not. The motor-control side — redistributing current across remaining phases within thermal limits, and observer-based fallback where the operating envelope permits — has to be co-designed with the battery-management system the drivetrain depends on.
Step 3: Validate every degraded mode with fault injection, not analysis alone. Exercise each fault on HIL under real-time dynamic scenarios, supplemented by power-HIL or dynamometer testing where the power stage must be exercised, measure the transitions into and out of each mode, and confirm diagnostic coverage and reaction time meet the safety target. For automotive programs this is ISO 26262 functional-safety validation backed by Automotive SPICE traceability; the analogous industrial standard is IEC 61508, and the two are not interchangeable. Simulation alone cannot demonstrate all real-system interactions or serve as the only validation evidence.
Real Trade-Offs
Hardware redundancy raises cost, weight, and part count, and each added component is one more thing to diagnose. The offset is removing the shutdown-only reaction that can create the vehicle-level hazard. This is a system-level call, most defensible on functions that directly influence vehicle motion, and it is settled by the ISO 26262 hazard and safety-goal analysis, not assumed.
A multi-phase power stage keeps torque available after a phase fault, but only if the topology was designed for it, and the remaining phases then carry more current — so the thermal and derating design must assume the faulted case.
Lockstep improves diagnostic coverage, but detection is not takeover: an independent controller or safety channel is needed where a function must keep running, which adds synchronization, state-sharing, and validation cost.
A buffered or independent low-voltage supply — a second DC/DC path or an LV buffer — keeps steering and braking powered through a converter fault, at the cost of a redundant path that must itself be maintained and monitored.
Operating in degraded modes or holding redundant paths can reduce peak efficiency and complicate thermal management. On functions where loss is recoverable — non-critical auxiliary loads — fail-safe remains the right and cheaper choice.
Typical Power-Electronics Safety Engineering Tasks
Safe-Transition Architecture
Hazard-driven decisions on degradation versus shutdown per fault, fault-tolerant time budgeting, and isolate-and-derate versus trip behavior for the power stage.
Redundant Power and Compute Design
Multi-phase and segmented power stages, buffered low-voltage paths, and independent monitor/second-controller architectures beyond lockstep detection.
Control-Firmware Safety Engineering
Mode-transition logic, observer-based sensor fallback within a validated envelope, and torque-capability coordination across inverter, brake, and vehicle-control units.
Fault-Injection and HIL Validation
Exercising every degraded mode under real-time dynamic scenarios and building the ISO 26262 evidence set.
Where This Sits
At this point the work is safety-architecture and validation, not another protection-circuit revision. The starting point is power-electronics architecture and design, where the safe-transition decision belongs. Signals that a project is here rather than in another test cycle: a bench that passes every injected fault while vehicle integration behaves differently, an inverter trip that removes torque in one uncoordinated step, or a degraded mode that is specified but never driven under load.
The logic that manages mode transitions and observer fallbacks lives in the MCU firmware and RTOS layer, and the cross-monitoring and system coordination sit in the ECU and embedded platform that ties the drivetrain to the rest of the vehicle.
FAQ
What is fail-operational in EV power electronics?
Why can fail-safe be insufficient for EV systems?
Does ISO 26262 require fail-operational design for every traction inverter?
Is a lockstep controller enough for fail-operational behavior?
Related Engineering Cases
- Power Inverter Architecture Design: 48 V MHEV inverter architecture: architecture selection, PLECS/Simulink, thermal, driving-profile verification. Relevant to power-stage and inverter architecture.
- Dual-MCU Railway BMU Architecture (SIL 2): Independent safety MCU with cross-monitoring. Adjacent proof of the monitoring pattern behind fail-operational control (railway SIL 2, not automotive).
- ASIL-C Current Sensor Design: Automotive current sensor to ASIL-C: functional-safety design and sensor integrity in the power path.