Why C-to-Rust Firmware Migrations Fail Without Architectural Redesign

Why C-to-Rust Firmware Migrations Fail Without Architectural Redesign

 

Quick Overview

Problem: A line-by-line Rust port of a working C firmware base tends to ship a wide unsafe surface that hides the same failure modes the migration was meant to close.

Common failure points: Global mutable state re-expressed behind static mut and unsafe; ISR/DMA sharing modelled without a designed ownership contract; vendor HALs wrapped rather than integrated at typed peripherals; toolchain decisions deferred past program-planning; hybrid C/Rust boundaries left informal.

Where it appears: Connected industrial controllers, automotive ECUs with growing connectivity, medical devices moving to networked operation, long-lifecycle IoT gateways, and a firmware base on a supported MCU target that has outgrown its original superloop.

Engineering focus: Ownership-driven module boundaries, ISR-safe shared state, DMA buffer lifetime contracts at the HAL, hybrid FFI surface design, and toolchain and certification planning matched to the target standard.

 

Illustrative scenario. The following example combines common embedded migration failure patterns and does not describe a specific Promwad client project.

The firmware built, flashed, and booted. Unit tests passed. What surprised the team was the shape of the field failures that arrived after the migration — not the migration itself.

A networked industrial controller on a Cortex-M7 MCU had gone through a full rewrite from a mature C base to Rust, driven by two chronic bug classes: a hard-to-reproduce DMA buffer race in the Ethernet path, and a slow drift toward "we cannot reason about this concurrency anymore" that came up at every quarterly review. Some categories improved. Others reappeared: intermittent corruption of a sensor reading tied to a specific IRQ preempting a specific task, and a class of DMA transfer errors that had never appeared in the C version.

Technically these were not Rust bugs. The team had shipped Rust code that walked around Rust: global state re-expressed as static mut behind unsafe blocks, ISR-to-task communication done with raw pointers because the ownership model "made it too hard," and the vendor HAL wrapped rather than re-architected. The compiler had done its job. The design had not.

Wrong Assumption

The assumption is reasonable: translate the working C firmware into Rust and get memory safety for free. That skips what actually delivers the guarantees. Rust's value lives in the ownership model, the borrow checker, and the Send / Sync rules that govern what may cross execution contexts — and those guarantees only bind when the system is structured to respect them. A codebase built around shared globals and informal module boundaries does not become memory-safe by acquiring Rust syntax; it becomes a Rust codebase with the same architecture and a growing unsafe surface hiding the same defects.

Why It Fails

Concurrency contracts that C leaves informal, Rust makes structural. In embedded C, shared state between the main loop, ISRs, and RTOS tasks is coordinated through convention: critical sections, hand-rolled mutexes, and the volatile qualifier for observable memory-mapped I/O access (noting that volatile is not a synchronization primitive and provides no atomicity or mutual exclusion). Rust makes the contract type-level: data crossing execution contexts must satisfy Send or Sync where the mechanism requires it. In RTIC, both #[shared] resources and #[local] resources moved from init must implement Send; only directly initialised task-local resources require neither.⁷ On multicore targets, a critical section does not by itself guarantee exclusivity between cores — additional synchronization primitives are required. The compiler will not accept a design that violates these rules, but it will accept an unsafe block that hides the violation. The same shared-state problem the event-driven active-object approach addresses at the RTOS layer reappears here one layer down.

DMA and ISR interaction is where the design pass earns its keep. DMA buffers must remain valid, aligned, and untouched by the CPU while the peripheral owns them. The Rust pattern is well-established: a HAL or driver models a DMA transfer as an ownership-holding value ("the transfer owns the buffer") and returns it only on completion. Enabling primitives like the embedded-dma trait crate exist for this — and it is worth being explicit that those traits are unsafe by design, because correctness depends on peripheral behaviour, cache coherency, and alignment guarantees the compiler cannot verify.⁸ Build the ownership pattern in and safe Rust prevents premature or aliased CPU-side buffer access; hardware, cache-coherency, memory-ordering, and driver-level unsafe defects still require target-level verification. Skip the pattern and cast a static array to a raw pointer inside unsafe, and you have paid the syntactic cost of Rust and kept the exact race.

The vendor HAL layer is where the ecosystem gap is real. C-based firmware sits on top of a mature stack: silicon vendor HALs (STMicroelectronics HAL/LL, NXP MCUXpresso SDK, Nordic nRF SDK, Renesas FSP), reference drivers, RTOS ports, certified middleware. Rust offers embedded-hal as a portable trait layer and per-vendor PACs generated from SVD files, plus HAL crates of varying maturity for popular Cortex-M and RISC-V targets. Coverage is strong for widely-used parts and thinner for less common silicon. Where a required peripheral has no idiomatic-Rust equivalent, the pragmatic answer is a defined FFI boundary to the vendor C code — a design decision, not an accident, with the ownership contract at the boundary spelled out. Our comparison of Zephyr, FreeRTOS, and ThreadX covers the RTOS-ecosystem trade-offs that shape which of these paths is available on a given program.

Toolchain and certification are separate from language choice. For safety-relevant firmware, mainline rustc is not the same object as a qualified compiler. Ferrocene, Ferrous Systems' downstream Rust toolchain, is TÜV SÜD-qualified for ISO 26262 up to ASIL D, IEC 61508 up to SIL 3, and IEC 62304 up to Class C, with a certified subset of core at ASIL B / SIL 2; it supports customer certification efforts toward IEC 61508 SIL 4 and DO-178C DAL C.¹ On Infineon AURIX (TC3x/TC4x), HighTec offers an alternative ASIL-D-qualified Rust compiler.² Where an ASIL-B or SIL-2 argument is on the table, mainline rustc cannot rely on an off-the-shelf pre-qualified compiler package unless tool confidence is addressed separately — which is a program planning task, not a compiler question. Where the argument is at the IEC 61508 level, a qualified-toolchain path is credible today.

Hidden System Complexity

source register / peripheral → PAC → HAL → driver → RTOS task or async executor → application layer → OTA / security surface → certification artifacts

A memory-safety guarantee at the application layer often depends on decisions three or four layers down. Ownership of a peripheral must be established at the PAC-to-HAL boundary and preserved upward. A HAL that exposes raw register pointers or permits multiple mutable owners of the same peripheral leaks ambiguity into every layer above it.

The interrupt path adds a second axis. RTIC makes the resource-to-context assignment part of the program's declared structure, so the compiler verifies who can access what;³ Embassy structures async tasks with an executor model that removes many concurrency errors by construction, though DMA, interrupt, and shared-state safety still live in the HAL, channel primitives, critical sections, and any unsafe code the driver ultimately calls.⁴

The OTA and security surface adds a third. Networked firmware ships with an update path, and that path is a high-value attack surface on the device. The memory-safety argument that motivates Rust in the network stack applies with more weight in the bootloader and update verifier, where a single parser bug can become a fleet-wide compromise — the class of failure that has motivated broader industry guidance on memory-safe languages.⁵ This is one candidate to assess for a targeted Rust module in a hybrid design; see secure OTA update pipelines and firmware integrity for the surface itself.

Failure Patterns

The following scenarios are illustrative patterns, not client project descriptions.

Scenario 1. A team migrates an industrial gateway from C to Rust and ships. The Rust code compiles clean under #![deny(unsafe_code)] at the crate root — a useful lint, though it does not cover dependencies and can be overridden inside modules — but vendor-adjacent code (driver wrappers, interrupt handlers, DMA setup) carries a large unsafe surface. Each block was opened to work around an ownership problem that would have required a redesign to solve properly. The original DMA race is unchanged; field crash rate is unchanged.

Scenario 2. An automotive supplier moves a body-control ECU from C to Rust to preempt future memory-safety pressure. The compiler is mainline rustc, so the ASIL-B safety argument cannot rely on an off-the-shelf pre-qualified compiler package unless tool confidence is addressed separately. The certification consultant flags this partway through, and the program either replans around a qualified toolchain (accepting its target and version constraints) or falls back to keeping the safety-relevant modules in MISRA-C and moving only the connectivity stack. The technical work was correct; the toolchain plan was not.

Scenario 3. A medical device team introduces Rust into the network-facing components of a Linux-based imaging platform. The FFI boundary is defined by "we'll pass pointers and check on the other side." Within a few sprints, a Rust-owned buffer is dropped early and a use-after-free or invalid memory access occurs in the C layer. The failure lands as a C crash, but the root cause is the missing ownership contract at the language boundary — the one thing a hybrid architecture cannot afford to leave informal.

Embedded Firmware Advisory — Architecture, Safety, and Modernisation Paths

Memory-safety failures in embedded firmware — ISR data races, DMA lifetime bugs, use-after-free in network stacks, unsafe creep in ports, and the toolchain gap between mainline Rust and qualified compilers — are architectural problems, not language problems. Promwad provides senior technical advisory on embedded architecture, platform selection, feasibility, safety risk, and lifecycle modernisation across MCU firmware and Linux-based embedded software in automotive, industrial, medical, and consumer connected devices. This article applies that assessment framework to a C-to-Rust migration decision.
 

Explore MCU Firmware Engineering →

MCU and Embedded Technology Experience

An Illustrative Composite: A Scoped Hybrid Migration

Illustrative composite scenario. The following is a composite of engineering patterns typical of modernisation reviews of connected industrial firmware. It is not a description of a specific client project and does not claim measured Promwad Rust delivery results. Its purpose is to illustrate the scoping and boundary-design decisions this article recommends.

Consider a networked industrial controller on a Cortex-M7 MCU whose C firmware has grown over several product generations. Chronic defects cluster in two places: intermittent DMA buffer issues in the Ethernet path, and memory-safety findings that fuzzing keeps re-surfacing in a TLV parser. The peripheral drivers, calibration code, and bootloader are stable, MISRA-audited, and tightly coupled to the vendor SDK — with no defect history.

The instinctive proposal is a full rewrite. A scoping review argues against it: rewriting stable, low-defect modules adds regression risk and re-verification effort against a limited near-term safety return, while the modules with the actual defects are a smaller share of the codebase.

A more defensible architecture keeps the peripheral drivers, bootloader, and calibration in C — unchanged — and moves the network stack, the TLV parser, and the configuration store to Rust behind a defined FFI boundary. Existing certification evidence for the unchanged modules may then be reused, subject to impact analysis and the applicable standard. The boundary is designed, not improvised: Rust owns buffers that cross it, C-side callbacks receive borrowed views with lifetimes documented on the Rust wrapper API (lifetimes do not travel across the C ABI itself), and every FFI function is covered by contract tests. On the Rust side, DMA descriptors are held by a driver-level transfer type that releases the buffer only after the transfer-complete interrupt fires; ISR-to-task communication goes through a typed channel rather than a raw pointer.

The point of the scenario: a targeted, hybrid C-to-Rust firmware migration matches the scope of the actual defect surface, keeps prior verification effort re-usable where modules are unchanged, and defines the language boundary as an engineering deliverable rather than an artifact of "we'll figure it out."
 

Scoped Hybrid Migration


Solution Approach

Step 1: Scope the migration to the defect surface, not the codebase. Pull field defects, static-analysis findings, and fuzzing reports for the last two release cycles, and cluster them by module. Modules with a documented memory-safety history — usually a connectivity stack, a parser, a serialization layer, or a security-critical component — are where a Rust migration returns value. Stable, well-tested, MISRA-audited peripheral and control code is usually where a rewrite adds risk without a proportional safety return. This scoping decision — treated as part of MCU firmware development discipline before any Rust code is written — separates migrations that ship from those that stall.

Step 2: Design the ownership model before touching a keyboard. For the target subsystem, answer for every resource: who owns it, when does ownership transfer, what execution context is allowed to touch it, and how does the type system express that. Peripherals become owned handles, not globals. DMA buffers become resources with an explicit loan-to-peripheral state carried by a driver-level transfer type. ISR-to-task communication goes through a typed primitive — an RTIC resource, an Embassy sync primitive, a lock-free channel — with the appropriate Send / Sync obligations. If this design is skipped, every hard question falls into unsafe.

Step 3: Define the hybrid boundary as an engineering artifact. Where C code stays — vendor SDK, legacy stable module, certified component — the FFI boundary is a contract with named owners, documented lifetimes on the Rust wrapper API, and CI enforcement of the surface. Lifetimes are not carried across the C ABI itself; the extern "C" functions traffic in raw pointers, lengths, and opaque handles.⁶ Validate ownership, pointer, length, nullability, aliasing, and callback-lifetime invariants through compile-fail tests, host-side harnesses, fuzzing, and C-side sanitizers where applicable — not by executing scenarios (early free, aliasing during a loan) at the boundary itself, which would exercise undefined behaviour.

Step 4: Plan the toolchain and certification path at day one. Decide whether the program will use a qualified Rust compiler (Ferrocene for ISO 26262 up to ASIL D, IEC 61508 up to SIL 3, IEC 62304 up to Class C on selected targets; HighTec for AURIX), a hybrid where safety-argued modules stay in MISRA-C, or a lower-assurance path where mainline rustc is acceptable. Each option constrains target selection, RTOS choice, and library surface. Re-verification of unchanged modules is normally an impact-analysis exercise rather than a full re-certification, but that judgment has to be made against the standard, not assumed. Deferring the toolchain decision can force significant rework when it is made under pressure at certification time. Our IEC 61508 functional-safety practice treats this as a program-planning input.

Real Trade-Offs

Full rewrite vs. hybrid architecture. A full rewrite gives a coherent codebase at the cost of longer schedules and re-verification of modules that had no defect history. A hybrid keeps the certified and stable code and adds an FFI surface to design, test, and maintain. On mid-size firmware bases with a well-defined defect cluster, hybrid can reduce rewrite scope and re-verification effort; on small greenfield systems, full-Rust is often cleaner.

Mainline rustc vs. qualified toolchain (Ferrocene, HighTec). Mainline Rust ships fast and tracks the latest language features. Qualified toolchains ship a certification artifact package and a slower cadence bounded by their supported targets. The choice is not about language capability but about which safety argument the program needs.

Async (Embassy) vs. RTIC-style static scheduling. Embassy suits connectivity-heavy firmware with an async I/O model and an active driver ecosystem. RTIC suits hard real-time work with static priority-based scheduling and compile-time resource-access verification. Mixing them is possible but adds complexity; picking the wrong one adds friction in every module.

MISRA-C hardening vs. Rust migration. For a C codebase whose defects cluster around a small parser or a network stack, a targeted MISRA-C hardening pass — tightening rule compliance, adding host-side sanitizer runs (ASan/UBSan) to CI (they generally run on host or simulation builds, not on bare-metal MCU targets), and expanding the fuzzing surface — may materially reduce the defect risk at a lower migration cost, although it cannot provide the same language-level guarantees a memory-safe language offers.⁵ Rust's value shows up when the memory-safety problem is systemic or when the networked attack surface makes the class of bug catastrophic. The RTOS landscape and MISRA alignment work covers how those trade-offs shift as safety and certification pressures rise.

Typical Firmware Engineering Tasks

Defect-Cluster Analysis and Migration Scoping

Field defect and static-analysis review to identify the modules where a Rust migration returns safety value and the modules where it adds regression risk without one.

Hybrid C/Rust FFI Boundary Engineering

Design, documentation, and contract testing of the language boundary in mixed-language firmware, including ownership handoff, safe-Rust wrapper APIs with expressed lifetimes, and CI enforcement of the surface.

Ownership and Concurrency Architecture Design

Redesign of peripheral ownership, DMA buffer lifetimes at the driver layer, ISR-to-task communication primitives, and RTOS task boundaries so type-level guarantees hold across execution contexts.

Toolchain, Safety Qualification, and Compliance Planning

Selection between mainline Rust, qualified toolchains (Ferrocene, HighTec), and hybrid MISRA-C paths against ISO 26262, IEC 61508, and IEC 62304 requirements, with target and library plans that match the chosen toolchain.

Qualifying Symptoms

  • A C firmware base has crossed the threshold where concurrency across the main loop, multiple ISRs, and RTOS tasks is no longer possible to reason about in code review.
  • Field defects and fuzzing findings cluster in the network stack, parsers, serializers, or a security-relevant component — memory-safety bugs that keep coming back in new variants.
  • A previous Rust proof-of-concept compiled and ran, but the team could not agree on how to handle ISRs, DMA, or the vendor HAL, and the effort stalled.
  • The product is networked or moving to networked operation, and the OTA and update-verification surface is a high-value attack path on the device.
  • A safety or security certification argument is on the horizon, and the toolchain decision has not been made against the specific standard the program will target.
  • A hybrid C/Rust codebase exists, but the language boundary is undocumented and defects on one side surface as unreproducible failures on the other.

At this point the work is architecture and toolchain planning, not another Rust proof-of-concept. In practice: defect-cluster analysis to bound the migration scope, an ownership and concurrency design pass before any translation begins, an FFI contract for the hybrid boundary, and a toolchain plan matched to the certification path.

For products where memory safety is loudest at the boot-and-update layer, the same discipline flows into secure OTA and firmware verification work. For Linux-based devices where the boundary sits at the kernel and driver layer, it shifts into Linux and Android kernel engineering, where the trade-offs — kernel modules in C, userspace daemons in a memory-safe language — repeat one layer up.

Where Promwad Fits

This class of problem shows up most in networked industrial controllers whose C base has outgrown its original superloop, automotive ECUs where connectivity is being added on top of a MISRA-audited safety layer, and long-lifecycle products where a language migration is being weighed against continued C hardening. Promwad's contribution to such programs is the architectural, safety-planning, and hybrid-integration assessment — the decisions that determine whether a migration returns its safety promise before any code is translated.

FAQ

Do we need a qualified Rust compiler, or is mainline Rust enough?

 

It depends on the safety argument. For internal or non-certified work, mainline Rust is fine. For ISO 26262 (ASIL A–D), IEC 61508 (SIL 2–3), or IEC 62304 (Class B/C), Ferrocene is TÜV SÜD-qualified up to ASIL D, SIL 3, and IEC 62304 Class C respectively, with a certified core subset at ASIL B / SIL 2, and supports customer efforts toward IEC 61508 SIL 4 and DO-178C DAL C.¹ On Infineon AURIX, HighTec offers an ASIL-D-qualified Rust compiler.² Mainline rustc can still be used in safety-argued programs if the team is prepared to build a tool-confidence argument itself, which is separate work.

 

Does the ownership model make ISR and DMA code harder to write than in C?

 

It pulls the design work forward. In C, an ISR that shares state with a task compiles and runs whether or not the sharing is correct; correctness is a review problem. In Rust, the type system pushes the decision up-front — sharing must be expressed as an RTIC resource, an Embassy sync primitive, a critical-section-guarded interior-mutability wrapper, or an equivalent. Day-to-day code can become comparable to C in effort once the patterns and tooling are established.

 

How do we design the hybrid C/Rust boundary so it does not become the source of new defects?

 

Treat the FFI surface as an engineering artifact. Name the owner of every buffer that crosses it, document lifetimes on the safe Rust wrapper API (the C ABI itself carries only pointers, lengths, and opaque handles),⁶ and validate the invariants — ownership, pointer, length, nullability, aliasing, callback lifetime — through compile-fail tests, host-side harnesses, fuzzing, and C-side sanitizers. Keep the surface narrow and enforce it in CI. The failure mode is almost always an informal boundary.

 

How much of our existing C code will we actually rewrite?

 

In a well-scoped migration, less than most teams initially assume. Peripheral drivers coupled to a vendor SDK, calibration code, and stable well-tested control loops usually stay in C — they have no defect history that Rust closes, and rewriting them is regression risk against a limited safety return. The value concentrates in networking, parsing, serialization, cryptographic and update-verification code, and modules with a documented history of memory-safety defects. The right proportion depends on the specific codebase; the scoping decision — not the rewrite itself — is what determines whether the migration ships.

 

Related Engineering Cases

Tell Us About Your Firmware Modernisation

We will assess the firmware architecture, defect concentration, toolchain constraints, and certification impact and recommend a practical modernisation path.

Tell us about your project

We’ll review it carefully and get back to you with the best technical approach.

All information you share stays private and secure — NDA available upon request.

Prefer direct email?
Write to info@promwad.com

Secured call with our expert in 24h
Secured call with our expert in 24h
Secured call with our expert in 24h
Plug-in model for your full-cycle R&D
Secured call with our expert in 24h
22 years of engineering expertise
Secured call with our expert in 24h
500+ projects for OEMs in EU & US
Secured call with our expert in 24h
MVP in 8–10 weeks — predictable delivery
Secured call with our expert in 24h
Featured at IBC, Embedded World, MWC