FOC Reference 06 Implementation

Software architecture

The control law is the part of a drive that is easiest to reason about and hardest to keep clean. It is a few hundred lines of arithmetic with no I/O in it — and in most firmware it ends up welded to a vendor SDK, a fixed-point library, an RTOS and a display stack, at which point you can no longer test it without a motor on the bench.

This page is one structure that avoids that, taken from a working drive rather than proposed in the abstract: an MCX-N947 running a Teknic PMSM, where a vendor-independent library was extracted step by step from an NXP SDK example while the motor kept running between every step.

Three layers, and a rule about which way they point

Three layers with the control law in the middle: the application points down into it, the port points up into it, and only the port names a chip. no vendor headers, no I/O calls the API implements it same contract the only vendor code app/ scheduling, UI, logging, tuning — the product core/ transforms, PI, modulator, observers, state machine port/<chip> measure, actuate, position, time — and nothing else vendor SDK drivers, clocks, register headers port/sim this site runs here
The control law sits in the middle of the stack, not at the bottom of it. A foundation is the thing a building stands on; the control law stands on nothing, which is why the arrows converge on it rather than passing through it. The port's arrow is the only one leaving the stack — that single line is everything the drive knows about a particular piece of silicon.

The dependency arrows all point inwards, and that single constraint is what buys everything else. The core cannot include a vendor header, so it cannot acquire a dependency on a chip. It has no I/O, so it can run anywhere — including in a unit test on your laptop, at ten thousand times real speed.

That is not a hypothetical benefit. This site is that test. Every figure here runs the same closed loop on a machine model, in a browser, because the control code is separable from the hardware it normally drives. The block marked port/sim is not an illustration; it is the block these pages execute.

On the bench The migration order matters more than the destination. Extracting a library in one pass gives you a firmware that has never run. Doing it in steps — maths first, then the port contract, then the current loop, then the state machine — with a bench run after each, means every regression has exactly one candidate cause. On the rig above, step 3 was checked by proving the new modulator agreed with the vendor’s to 0.009% before it was allowed to drive anything.

The port contract is the whole design

Everything platform-specific reduces to a handful of operations:

OperationWhat the core needs
measuretwo phase currents and the DC-bus voltage, from this period
actuatethree duty cycles, applied coherently
positionencoder count and index, when one exists
timethe control period, as a number

Nothing else. No init, no clocks, no interrupt priorities — those are the app’s business, and keeping them out of the contract is what stops the port layer growing into a second SDK.

Write that contract down before porting anything. It is also the place where a second implementation costs nothing: a simulated port that feeds the core from a machine model, which is how the control law gets exercised without a bench.

Two clocks, and which one owns what

There are two clocks in a drive and confusing them is a common source of mysterious behaviour.

Fast, in the ISR, once per PWM period. Read currents, Clarke, Park, two current PIs, decoupling, inverse Park, modulate, write duties. Nothing else. No allocation, no logging, no waiting. On the rig this is about 12 µs of a 62.5 µs period — a measured number on that hardware, not a model output.

Slow, in a task or a loop, typically 1–10 ms. Speed and position loops, ramps, fault policy, the UI, telemetry. These run at a rate the mechanical system can actually respond to, and none of them belong in a 16 kHz interrupt.

The two communicate through plain shared state — a small struct of commands one way and measurements the other. If that boundary needs a mutex, the split is in the wrong place.

The same control law, scheduled two ways

Here is the part that is usually presented as a matter of taste and is not. The fast loop is identical under an RTOS and under a bare-metal superloop: it is an interrupt in both, it preempts everything in both, and it contains exactly the same arithmetic. Nothing about field-oriented control needs a scheduler.

What differs is the layer underneath, and the difference is not stylistic.

Under an RTOS: the control ISR at the top, a periodic control task below it, the UI below that, and one shared command block between them. measurements references read only ADC complete hardware, every 62.5 µs control ISR current loop only — no allocation, no blocking command block plain struct, one writer per field control task speed loop, ramps, faults — 1 kHz timer UI and telemetry lowest priority, allowed to run long
Three priority levels. The control interrupt is not a task and never becomes one. Under it, a periodic task carries the work that may be a millisecond late; under that, everything that may be a second late. The command block is a plain struct with one writer per field — the moment it needs a lock, the split between the rates is in the wrong place.
Bare metal: the same control ISR, with one main loop underneath carrying the slow control work and all the background work at the same priority. measurements references then this round again ADC complete hardware, every 62.5 µs control ISR the only interrupt that matters — identical code command block plain struct, one writer per field main loop slow control work, when the loop gets to it everything else UI, telemetry, comms — same priority
Drawn on the same frame, so the difference is the only thing that moves. The top half is identical. Underneath there is one priority level instead of two, so the slow control work shares it with everything else — and the slow loop's period stops being a number anyone configured. It becomes however long a trip round the loop happens to take.

That last sentence is the whole trade, and it is measurable rather than arguable. Give both models the same interrupt, the same slow work and the same background chunk, and watch what each delivers:

The fast loop
The slow loop
Top: one PWM period at its true scale, the same under both models — the interrupt takes its share before anything else runs. Bottom: several milliseconds of the same drive under each model, with a marker at every slow-loop release. Drag Background up and the FreeRTOS trace does not move, while the bare-metal releases spread out in proportion. Drag it down far enough and the superloop overtakes the requested rate instead — a superloop has no rate at all, only a loop time.

At the defaults — 16 kHz, a 12 µs interrupt, a 1 kHz slow loop and a 3 ms background chunk — the interrupt takes 19.2% of the CPU under both models. The RTOS delivers the 1 kHz. The superloop delivers about 266 Hz, because one trip round its loop is 3.04 ms of work stretched by the interrupt into 3.76 ms of wall clock. Nothing is broken. The rate was simply never being scheduled.

On the bench This is not a thought experiment. On the rig behind this page the speed reference generator lived in the user-interface task at a nominal 10 Hz. Measured, that task ran at 2–3 Hz under chart-rendering load, so a commanded 500 rpm/s ramp was delivered at about 80 rpm/s — the drive felt sluggish for reasons that had nothing to do with the motor, the gains or the machine.

The fix was not a faster task. It was moving reference generation and the outer loops into the control interrupt, where the rate is a property of the hardware rather than of how busy the display is. That is worth considering before reaching for an RTOS: promoting the work is often cheaper than promoting the scheduler.

So: use an RTOS when something else in the product genuinely needs one — a network stack, a filesystem, a graphics library that expects to block. Do not adopt one to make the control loop deterministic. It already was, and it was determined by an interrupt vector rather than by a scheduler.

One owner of the inverter

The single hardest bug on the rig behind this page came from breaking that rule: two pieces of code wrote the duty registers in the same interrupt, and the hardware’s reload rule decided which write survived. The symptom was physics that could not happen; the cause was ownership, not arithmetic.

So make it structural. Exactly one function writes duties. Everything that wants to influence the inverter returns a value to that function instead of reaching for the peripheral. A drive with two writers cannot be made correct by ordering them carefully — see the peripherals page, where the mechanism is drawn out period by period and you can watch the hardware, not your code, pick the winner.

A state machine that owns the transitions

Modes are where drives break, because the interesting failures live in the transitions rather than the states: aligning while the ADC is still calibrating, handing over from open loop to closed loop at the wrong angle, opening the contactor at speed and regenerating into the bus.

The drive's mode machine: BOOT to READY to CALIBRATE to ALIGN to RUN, with FAULT reachable from everywhere and a ramped path back to READY. init done start offsets angle known stop, ramped no hardware guard trip cleared BOOT clocks, peripherals, self-test READY outputs disabled, waiting CALIBRATE ADC offsets, bridge idle ALIGN rotor to a known angle FAULT outputs safe, cause latched RUN the current loop is closed
Only two paths into FAULT are drawn; in a real supervisor every state exits to it, and every such exit tapers the current before modulation is removed. Two arrows on this diagram are labelled with the bugs they fix: offsets, because the supervisor once started before calibration finished and injected current into its own zero reading, and stop, ramped, because a step command to zero at speed regenerated into the link and tripped over-voltage.

A small explicit machine, with every transition in one file, is worth more than the sophistication of any individual mode. The two failures named in the caption were both transition bugs rather than control bugs, and neither was visible in any single state’s code.

Guards, and a witness

Two things are worth building in from the start rather than adding after an incident.

Guards. A drive with an estimator can poison itself: on the rig, a stalled sensorless startup produced a NaN in the observer, which propagated into the voltage request and silently killed the drive until reboot. A check for non-finite values in the fast loop, which scrubs the PI accumulators and forces a fault, costs a few instructions and turns a mystery into a fault code.

A witness. Keep an independent measurement that the control law does not use. An encoder in a sensorless drive is the obvious one — when the observer reports a healthy 435 rad/s and the encoder count is frozen, you know immediately which one is lying. Ground truth you are not using is not waste; it is the only thing that can contradict a confident estimator.

What to take away

  • Point every dependency inwards, and the control law becomes testable without hardware.
  • Keep the port contract to measure, actuate, position, time.
  • Fast loop does the current loop and nothing else; everything slower goes in a task or a loop.
  • The current loop does not need an RTOS. What needs one is the slow work, and only if something is going to make the loop long — under a superloop, the slow period is the loop time, not a setting.
  • One writer for the duty registers, enforced structurally.
  • Put the transitions in one place, guard against non-finite values, and keep a measurement you do not use.