Single-phase smart-meter SoC · study guide · 2026-09-29
AFE-to-MAC datapath
One sample of voltage, current and neutral current travels from the Dolphin AFE to the CPU in seven stages. Each stage below starts with a plain-words explanation, then gives the exact widths, addresses, handshakes and reasons. Every number is worked out on the page.
Basis of this guide
What this guide is based on
Tags: spec comes from the Dolphin specification, assumed is our own working choice, not yet confirmed, and open is unresolved (see the last section).
| Parameter | Value | Status |
|---|---|---|
| Scope | Single phase | assumed |
| Channels | V line voltage · I phase current · T neutral (tamper) current | assumed |
| AFE | Dolphin Metro-PM-MFE, serial DSP mode, 24-bit two's complement | spec |
| Sample rate | 4,000 sps per channel, fixed. MCLK 4.096 MHz ÷ 4 = 1.024 MHz sync rate ÷ 256 = 4 kSps | spec |
| Stored word | 32-bit, sign-extended from 24-bit | assumed |
| Path | AFE → capture → FIFO → DMA → SRAM ping-pong → MAC reads SRAM as AXI master | assumed |
| MAC clock | 100 MHz, clock-gated, one shared 32×32→64 multiplier | assumed |
| CPU clock | Up to 200 MHz; 100 MHz = PLL 200 MHz ÷ 2, synchronous | open O3 |
| Buffer | 2 halves × 3 channels × 80 samples × 4 B = 1,920 B | assumed |
| HW / SW split | MAC hardware: every per-sample operation. Firmware: per-window operations, 5 per second | assumed |
Overview
The path in seven stages
The three channels stay in separate lanes (colour = channel) until they meet in the MAC. Pick a stage, or press Play. Every arrow shows its width and its rate. Change the Variables below the diagram and the labels, the tables and the numbers in every stage update.
What one word looks like at this stage (bit widths)
Variables
Assumed values are preset. Change any field to see a what-if; fields that differ from the assumed values are tagged. Nothing here changes the spec.
Results with the assumed values are: 48 kB/s into SRAM, 96 kB/s at the SRAM port (0.024 % of the bus), 57-bit sums, 147× headroom. The 25 result words per window are counted from the register map (my count). The MAC schedule is defined for 3 channels only, so the 7-channel option changes traffic and buffer sizes but reuses the same MAC cycles.
Stage 1 of 7 · AFE
AFE output: three serial words per sample instant
In plain words
The AFE (analog front end) is a separate Dolphin chip. It measures the voltage (V), the phase current (I) and the neutral current (T), and turns each into a 24-bit number. Every 250 µs there is a new number for each of the three.
It sends each number one bit at a time. Each channel uses 3 wires that carry one bit each: a bit clock (SCLK), a "new word starts" pulse (SSYNC) and the data (SDATA). Three channels make 9 wires. One bit at a time keeps the wire count small.
How fast SCLK runs is decided by the AFE and the Dolphin spec gives no number. It has to be at least 100 kHz (25 bits must fit in 250 µs) and cannot exceed the AFE's own 4.096 MHz clock.
"72 bits" is 3 channels × 24 bits. No 72-bit bus exists in the SoC: each channel has its own 3-wire serial bus, and each bus carries one 24-bit word per sample instant.
What leaves the AFE
| Item | Value | Source |
|---|---|---|
| Channels | V unity-gain buffer, DO = 1006633 · Vin · 2 · I PGA ×4/8/16/32 · T PGA ×4/8/16/32 | Dolphin §1.1, §3.6 |
| Modulator / sync rate | 1.024 MHz = MCLK 4.096 MHz ÷ 4 | §5.3 p.57 |
| Decimation | ÷256 → 4,000 sps (MCLK ÷ 1024 overall) | §3.4 |
| Word | 24-bit two's complement, range −8,388,608 … +8,388,607 (0x800000 … 0x7FFFFF) | §3.6 |
| Real resolution | Bits [2:0] are tied to 0 → code step 8, about 21 effective bits | §5.4.2 p.59 |
| Code → volts | current: Vin = code / (1006633 × gain) | §3 |
| Sampling instant | Simultaneous: all three channels share one MCLK-synchronous sample edge | §4.2.2 |
| Interface | DSP mode: MFE_SCLK, MFE_SSYNC, MFE_SDATA per channel (×3 buses), plus IRQ_V/I/T | §2.2, §5.4 |
| Bits per instant | 3 × 24 = 72 bits → 288 kbit/s at 4,000 sps | |
| SCLK rate | Set by the Dolphin AFE. The spec gives no value (§5.4 says only that the signals are synchronous to MCLK, 4.096 MHz). What the design can prove without it: SCLK must be at least 25 bits ÷ 250 µs = 100 kHz, and at most MCLK = 4.096 MHz. One 24-bit word then takes between 5.9 µs and 240 µs. The page uses 2.048 MHz (11.7 µs) only as an example value | open O2 |
| Bit order | §5.4.2 text says LSB-first, Fig 5.4 shows MSB-first. The capture RTL must make this a parameter until Dolphin answers. | open O1 |
Frame timing on one bus
Inside the AFE, before the pins
| Block | Job | Parameter |
|---|---|---|
| Sensor + anti-alias RC | Shunt, CT or divider drives a differential pair | RAA 1 kΩ, CAA 10 nF |
| PGA | Scales the sensor level to ADC full scale | I, T: ×4 … ×32 · V: fixed |
| ΔΣ modulator | 1-bit stream at 1.024 MHz carrying the signal in its density | order and OSR not stated |
| Decimation filter | Removes out-of-band noise, drops the rate to Fs | ÷256 at 4 kSps |
| HPF | Removes DC and offset | corner not specified |
| Calibration | Gain and offset trim per channel | GAIN_ERR_x, OFFSET_x |
| Phase shifter | Aligns V against I after sensor and filter delays | PSH_x: up to +17.9° at 50 Hz, step 0.0176° |
| Serialiser | Shifts the finished 24-bit word out | see frame timing above |
Why this matters downstream: V and I are sampled at the same instant and phase-aligned inside the AFE. The SoC never re-aligns them; it only has to keep them index-aligned (sample n of V pairs with sample n of I). Every buffer choice below protects that.
Stage 2 of 7 · Capture
Capture: three wires become three 32-bit words
In plain words
The capture block waits for the "new word" pulse, then shifts each arriving bit into a 24-bit register. After 24 clock ticks the number is complete. It then copies the top bit (the sign bit) into 8 new bits on the left, which makes a 32-bit number.
Why widen to 32 bits: the bus, the memory and the MAC all work in 32-bit words. Adding zeros on the left would turn a negative number into a huge positive one. Copying the sign bit keeps the value the same. This is called sign extension. Try it:
Try -5, then 5, then 8388607 (the biggest 24-bit value). The 8 grey bits on the left are always copies of the red sign bit.
A dedicated capture slice per channel turns the wiggling pins into one sign-extended 32-bit word. The three slices run side by side; their words complete within the same few clocks.
Mechanism, in order
| # | Step | Detail |
|---|---|---|
| 1 | Synchronise | SCLK, SSYNC and SDATA of each bus pass through 2 flip-flops into the 100 MHz domain. 3 buses × 3 signals = 9 synchroniser cells. Whatever Dolphin picks, SCLK ≤ 4.096 MHz, so the 100 MHz clock oversamples it at least 24×. |
| 2 | Find the SCLK edge | sclk_rise = sclk_s & ~sclk_s_d, one 100 MHz clock wide. |
| 3 | Arm on SSYNC | SSYNC seen → bit counter = 0, shifter armed. Data starts one SCLK later. |
| 4 | Shift | On each sclk_rise while armed: shift sdata_s into the 24-bit register, counter + 1. |
| 5 | Word done | Counter = 24 → word_ready pulse, d[23:0] latched. |
| 6 | Sign-extend | word32 = {{8{d[23]}}, d[23:0]} |
| 7 | Push | Write word32 into that channel's FIFO. Three slices push independently. |
Watch one 24-bit sample arrive
Press Play. On every clock tick the AFE sends one bit and the capture block drops it into a 24-slot register. After 24 ticks the number is complete, and the top bit is copied to make a 32-bit word.
1 · Bits arriving on the wire (first bit on the left)
2 · The 24-slot register (slot numbers below)
3 · The 32-bit word that goes to the FIFO
Advanced: sample value in hex, and what if the bit order is wrong?
The Dolphin spec contradicts itself on bit order (open item O1). Set the two menus to different orders and press Play: the number comes out bit-reversed and nothing in the datapath raises an error.
Why this matters: a wrong bit order gives a bit-reversed number with no error flag anywhere in the datapath. That is why O1 must be closed with Dolphin before RTL freeze, and why the receive order should be a register bit until then.
Sign extension
| 24-bit code | Signed value | 32-bit word stored | Name |
|---|---|---|---|
| 0x7FFFF8 | +8,388,600 | 0x007FFFF8 | highest value the AFE can send (low 3 bits are always 0) |
| 0x000008 | +8 | 0x00000008 | smallest positive step shown here |
| 0xFFFFF8 | −8 | 0xFFFFFFF8 | small negative |
| 0x800000 | −8,388,608 | 0xFF800000 | lowest value (most negative) |
| Choice | Reason |
|---|---|
| 32-bit word, not 24 | SRAM and AXI are 32-bit; a 24-bit packing would need 3-byte addressing and unaligned reads. |
| Sign-extend, not zero-pad | The MAC multiplier is signed 32×32. The word is already a valid int32_t, so neither the MAC nor the C code masks or extends anything. Zero-padding would turn −8 into +16,777,208. |
| Cost | 8 extra bits per word = 25 % storage overhead = 48 kB/s instead of 36 kB/s. Negligible. |
Stage 3 of 7 · FIFO
FIFO: one small queue per channel
In plain words
The 32-bit word goes into a FIFO (first in, first out). Each channel has its own FIFO holding 4 words, and words leave in the order they came in.
Why a FIFO: the capture block produces a word when the AFE says so, and the DMA collects it when the bus is free. Those moments do not line up, so the FIFO holds the word in between. Four words at 4,000 samples per second is 1 ms of slack.
Why one per channel: if V, I and T shared one FIFO and one word were ever lost, everything after it would shift by one place and V would land in I's slot. Separate FIFOs keep a loss inside one channel. Three FIFOs is our assumed choice; one shared FIFO is the alternative.
| Parameter | Value | Status |
|---|---|---|
| Count | 3: FIFO_V, FIFO_I, FIFO_T | assumed |
| Geometry | 32 bits wide × 4 words deep, 16 B each. Width follows the 32-bit word. Depth 4 is a margin choice with no published benchmark: keep it a parameter (DEPTH) and prove it with the stall test. | assumed |
| Clocking | Synchronous. Write side and read side are both in the 100 MHz domain; the clock-domain crossing was done by the 9 synchronisers in capture. | assumed |
| Arrival rate | 1 word per channel per 250 µs (4,000 words/s each, 12,000 words/s total) | |
| Slack | 4 words = 1 ms per channel. The steady-state need is under 1 word. | assumed |
| DMA request | Asserted while the FIFO is not empty (one request line per stream) | assumed |
| Flags | Per FIFO: empty, sticky overflow | assumed |
| Overflow action | Not defined by Dolphin. Proposal: drop the new word, set the sticky flag, raise the error interrupt. | open O8 |
Stage 4 of 7 · DMA and SRAM
DMA writes the words into SRAM ping-pong halves
In plain words
The DMA (direct memory access) is a small hardware courier. It takes each word out of a FIFO and writes it into SRAM, then moves on 4 bytes to the next slot. The CPU does none of this.
SRAM is split into two halves of 80 samples each. The DMA writes into one half, and the MAC reads the other half, which is already full. When the DMA finishes its half, the two switch jobs. This is called ping-pong. Filling a half takes 80 ÷ 4,000 = 20 ms.
The MAC is fast. When a half is full, the DMA sends a half_done event. The MAC then reads all 80 samples of V, I and T and goes idle. The DMA needs about 20 ms to fill the other half, so the MAC always finishes long before that half is full.
What switches is the address. The DMA moves its write address to the other half by itself, so it never stops. The MAC is not moved: it is started on the half that just filled. half_done for half 0 means "read the half 0 addresses", then half 1, then half 0 again.
| Time | DMA writes | MAC reads |
|---|---|---|
| 0 to 20 ms | half 0 | idle |
| 20 ms | switches to half 1, sends half_done for half 0 | reads half 0 |
| 20 to 40 ms | half 1 | idle |
| 40 ms | switches back to half 0, sends half_done for half 1 | reads half 1 |
If the MAC were still busy when the next half_done arrived, the BLK_OVR flag would be set so the problem is visible. That cannot happen while the MAC has 20 ms and needs only a tiny fraction of it.
DMA channels
| Field | ch0 (V) | ch1 (I) | ch2 (T) |
|---|---|---|---|
| Source | FIFO_V data reg, fixed | FIFO_I data reg, fixed | FIFO_T data reg, fixed |
| Destination | V region, +4 per word | I region, +4 per word | T region, +4 per word |
| Width | 32 bit | 32 bit | 32 bit |
| Count | 80 words per half | 80 | 80 |
| Mode | Ping-pong, auto-swap | same | same |
| Trigger | FIFO_V not empty | FIFO_I not empty | FIFO_T not empty |
| Aspect | Value | Status |
|---|---|---|
| Per trigger | 1 AXI read of the FIFO register and 1 AXI write to SRAM, 4 bytes each: ARLEN=0, ARSIZE=2, AWLEN=0, AWSIZE=2, WSTRB=4'b1111 | |
| Interrupt granularity | One event per 80 words, not per word | |
| Latency budget | Request to bus transaction ≤ 15 µs under full contention (specified at 32 kSps; at 4 kSps the sample period is 250 µs). A requirement, not yet shown to be met by the bus design. | assumed |
| Priority | Metering A above comms B above memory copy C, pre-empted at beat boundaries | |
| Overrun | Per-channel flag if the write pointer laps unread data | |
| Protection | Metering channels software-lockable | |
| DMA traffic | 3 × 4 B × 4,000 = 48 kB/s written to SRAM, 48 kB/s read from FIFOs | computed |
Wait for all three channels. V, I and T each have their own DMA channel, and they finish their 80th word a few microseconds apart. Say V finishes first and T last. If the MAC started as soon as V was done, it would reach the end of the T data before the DMA had written the last T word, and it would read an old value. So the MAC must wait until all three channels have finished (V and I and T), and only then start. The design notes say "half-done starts the MAC" but do not say this wait is built in, so it is open item O6.
SRAM layout
Planar: each channel is one contiguous array of 80 words per half. That matches three independent DMA channels with one incrementing destination each, and one source-address register per channel in the MAC. A whole window of a single channel is also contiguous for later waveform or AI use. Base address 0x2000_0000 is assumed (the SoC memory map is not defined, O4).
| Region | Address range | Bytes (one channel, one half) | MAC register |
|---|---|---|---|
| V half 0 | 0x2000_0000 – 0x2000_013F | 320 | MAC_SRC_V0 |
| I half 0 | 0x2000_0140 – 0x2000_027F | 320 | MAC_SRC_I0 |
| T half 0 | 0x2000_0280 – 0x2000_03BF | 320 | MAC_SRC_T0 |
| V half 1 | 0x2000_03C0 – 0x2000_04FF | 320 | MAC_SRC_V1 |
| I half 1 | 0x2000_0500 – 0x2000_063F | 320 | MAC_SRC_I1 |
| T half 1 | 0x2000_0640 – 0x2000_077F | 320 | MAC_SRC_T1 |
| Total, 6 rows | 1,920 | MAC_BLK_LEN = 80 | |
One row above is one channel in one half: 80 samples × 4 B = 320 B. A channel has two rows (half 0 and half 1), so one channel uses 2 × 320 = 640 B in total. One half holds three channels: 3 × 320 = 960 B (this is the 0x3C0 step between halves). Both halves: 1,920 B.
Address of sample n: base + half × 0x3C0 + channel × 0x140 + n × 4, with channel 0 = V, 1 = I, 2 = T and n = 0 … 79. Memory is little-endian: a stored word holds d[7:0] at the lowest address, then d[15:8], d[23:16], and last the sign byte 0x00 or 0xFF.
Address calculator
Ping-pong timing
| Event | Effect |
|---|---|
| Half done (all 3 channels) | MAC starts on that half (MAC_CTRL.AUTO = 1). DMA continues into the other half. |
| Half done while MAC still busy | MAC_STATUS.BLK_OVR = 1, sticky. The MAC needs a tiny fraction of the 20 ms half, so this only triggers on a fault. |
| Why two halves | The MAC is never reading a slot the DMA is writing, without any lock or copy. |
Stage 5 of 7 · MAC read side
The MAC reads SRAM as an AXI master
In plain words
The MAC does not wait to be handed data. It sends its own read request: an address goes out on the bus and SRAM answers with one 32-bit word. That makes the MAC the bus master.
For every sample it makes 3 reads (V, I, T). The data is read straight from the buffer where the DMA left it, so it is never copied a second time.
The MAC generates its own read addresses. Nothing pushes operands to it: once started it walks the half sample by sample, three single-beat reads per sample (V, I, T).
Sequence per block
| # | Step | Detail |
|---|---|---|
| 1 | Start | All-channel half-done event with MAC_CTRL.AUTO = 1, or MAC_CMD.START0/START1 written by the CPU. Clock gate opens. |
| 2 | Pick the half | The half number selects MAC_SRC_V/I/T[h], three 32-bit base addresses. |
| 3 | Address | Sample counter n = 0 … MAC_BLK_LEN−1 (80). ARADDR = SRC_ch[h] + 4·n for each of V, I, T. |
| 4 | Read | Issue AR_V(n), AR_I(n), AR_T(n) back to back; accept three R beats. |
| 5 | Compute | The three words are the inputs for the seven sums (next stage). |
| 6 | Prefetch | Reads for sample n+1 are issued while sample n is still being processed, so the multiplier does not wait for the bus. |
| 7 | Finish | n reaches 80 → BUSY clears, clock gate closes. |
ARADDR = V[5] 0x2000_0014, I[5] 0x2000_0154, T[5] 0x2000_0294. A request is accepted on the clock where ARVALID and ARREADY are both high.AXI signals used
| Channel | Signal | Width | Value from the MAC |
|---|---|---|---|
| AR | ARADDR | 32 | byte address, word-aligned (bits [1:0] = 00) |
| AR | ARLEN | 8 | 0 → one beat |
| AR | ARSIZE | 3 | 3'b010 → 4 bytes |
| AR | ARBURST | 2 | 2'b01 INCR (no effect at length 1) |
| AR | ARID | 2 | assumed 0 = V, 1 = I, 2 = T, so returns can be matched |
| AR | ARVALID / ARREADY | 1 / 1 | Master drives VALID and holds address until READY; VALID must not wait for READY |
| R | RDATA | 32 | the stored sign-extended word |
| R | RRESP | 2 | expected 2'b00 OKAY. Handling of any other value is not defined (O8). |
| R | RLAST | 1 | 1 on the only beat |
| R | RVALID / RREADY | 1 / 1 | MAC holds RREADY high whenever it expects data |
| Parameter | Value |
|---|---|
| Read traffic | 3 × 4 B × 4,000 = 48 kB/s |
| SRAM traffic total | DMA write 48 kB/s + MAC read 48 kB/s = 96 kB/s = 0.024 % of a 32-bit 100 MHz bus (400 MB/s) |
| Read latency | 2 clocks after accept assumed (synchronous SRAM); hidden by prefetch |
| Clock | MAC, DMA, capture and the bus in this path are all 100 MHz, synchronous. No clock-domain bridge on the sample path. |
| Memory port | The buffers could sit in tightly-coupled memory or in SRAM read over AXI. Which port and latency apply is open (O7). |
Why the MAC is the master
Suppose the CPU fed the MAC instead. Each operand would be read from SRAM by the CPU and written to the MAC, so it would cross the interconnect twice, and the CPU would run a loop 4,000 times a second. With the MAC as master, each operand is read once, by the MAC itself. The CPU has nothing to do per sample and can sleep between windows.
Stage 6 of 7 · MAC compute
Seven accumulations per sample
In plain words
The MAC (multiply and accumulate) is hardware that runs by itself. For every sample it multiplies and adds into 7 running totals: V×V, I×I, T×T, V×I, V×T, and two more that use a copy of the voltage delayed by a quarter wave (used for reactive power).
The CPU only sets it up (where the buffers are, how long) and switches it on. The multiplying is wired hardware, not program instructions. The totals are 64 bits wide because a product is up to 46 bits and up to 889 are added in one window, which needs 57 bits.
The MAC has 25,000 clocks per sample (100 MHz ÷ 4,000 sps) and needs only 8 multiplies, so it is idle almost all the time and its clock is gated off between blocks.
Inputs and internal state
| Item | Width | Meaning |
|---|---|---|
| v, i, t | 32 signed | the three words just read |
| Delay line | 32 × 32 | circular register file, one v written per sample |
| D_INT | 5 | MAC_DLY[4:0]: whole-sample part of the 90° delay |
| D_FRAC | 16 | MAC_DLY[31:16]: fractional part, Q16 (0x8000 = 0.5) |
| x0, x1 | 32 signed | two adjacent delay-line taps around the delay |
| v90 | 32 signed | x0 + (((x1 − x0) × D_FRAC) >>> 16), voltage delayed by a quarter cycle |
| Quantity | Value |
|---|---|
| Delay needed | Fs / (4 f) = 1000 / f samples: 20.000 at 50 Hz, 22.2 at 45 Hz (worst case, so the line is 32 deep) |
| Why 90° | Reactive power is the average of v(t − T/4) × i(t). A quarter period is an exact integer only at 50.000 Hz, so the fraction is interpolated between two neighbouring samples. |
| Who sets the delay | Firmware, from the measured frequency, once per window. No hardware divider. |
The seven accumulators
| k | Name | Sum | Becomes | Register (assumed order) |
|---|---|---|---|---|
| 0 | Σv² | v × v | V rms | ACC_LO/HI(0) 0x040 / 0x044 |
| 1 | Σi² | i × i | I rms | 0x048 / 0x04C |
| 2 | Σt² | t × t | T (neutral) rms | 0x050 / 0x054 |
| 3 | Σv·i | v × i | active power, phase | 0x058 / 0x05C |
| 4 | Σv·t | v × t | active power, neutral | 0x060 / 0x064 |
| 5 | Σv90·i | v90 × i | reactive power, phase | 0x068 / 0x06C |
| 6 | Σv90·t | v90 × t | reactive power, neutral | 0x070 / 0x074 |
Accumulator width
| Step | Value |
|---|---|
| Largest input | |code| = 2²³ (at 0x800000) |
| Largest product | 2²³ × 2²³ = 2⁴⁶ → 48-bit signed |
| Largest window | 889 samples (10 cycles at 45 Hz) |
| Largest sum | 889 × 2⁴⁶ = 6.26 × 10¹⁶ ≈ 2⁵⁵·⁸ → 57 bits signed |
| Register | 64-bit signed → 2⁶³ / 6.26 × 10¹⁶ = 147× headroom (2¹⁷ = 131,072 samples, 32.8 s, before any overflow) |
| Why 64, not 57 | The multiplier produces 64 bits, and a 64-bit value reads cleanly as two 32-bit words over APB. |
| If it ever overflows | MAC_STATUS.ACC_OVF = 1, sticky |
Cost of the per-sample work
| Quantity | Calculation | Value |
|---|---|---|
| Clocks available per sample | 100 MHz ÷ 4,000 | 25,000 |
| Multiplies per sample | 7 sums + 1 for v90 | 8 |
| Multipliers | 1 shared instead of 7 parallel | ≈ 1/7 area |
8 multiplies against 25,000 available clocks: timing is not a design constraint. How many clocks the MAC actually needs depends on the multiplier and control logic chosen, and is not fixed here.
Stage 7 of 7 · Window close
Windows close on zero crossings, then the CPU takes over
In plain words
A window is the group of samples that makes one result: exactly 10 mains waves. At 50 Hz that is 10 × 20 ms = 200 ms, or 800 samples. It is not the same as a buffer half: a half is 80 samples of memory, so one window spans 10 halves.
At the end of a window the MAC copies its 7 totals into shadow registers, raises a flag and starts the next window straight away. The CPU reads a 64-bit total as two 32-bit halves, and an add in between would give half-old, half-new data. The frozen shadow copy avoids that.
The CPU then does the last steps from the shadow totals: square root for RMS, scaling to volts and amps, energy and events. That happens 5 times a second. The window is set to end on a voltage zero-crossing so it always covers whole waves.
10 here means 10 mains cycles (10 waves of the 50 Hz supply = 200 ms), not 10 MAC clock ticks. See Numbers.
Windows are not the same thing as buffer halves. The DMA halves are fixed at 80 samples; a window is 10 mains cycles found by the MAC from the voltage zero crossings, about 800 samples, spanning roughly ten halves. Averaging over whole cycles cancels the 2f ripple in v·i and v² that a fixed block would leave when the line is not exactly 50.000 Hz.
Window configuration
| Field | Register | Value | Derivation |
|---|---|---|---|
| Cycles per window | MAC_WIN_CFG[7:0] | 10 | 10 cycles = 200 ms at 50 Hz |
| Max samples | MAC_WIN_CFG[23:8] | 889 | 10 × 4,000 ÷ 45 + 1 |
| Min crossing gap | MAC_WIN_CFG[31:24] | 57 | 4,000 ÷ 70 Hz, rejects double crossings |
| Re-arm hysteresis | MAC_ZC_HYST[23:0] | codes | a crossing re-arms only after |v| exceeds it |
Behaviour rules
| Event | Action |
|---|---|
| Positive crossing, armed, ≥ min gap since the last | No window open: open one and latch ZC0. Otherwise count it; on the 10th, close (FREQ_OK = 1, latch ZC1) and open the next (ZC0 = this crossing). |
| First crossing after start or after a forced close | Close whatever accumulated with FREQ_OK = 0, then open a window. |
| N reaches max samples (889) | Close with FORCED = 1, FREQ_OK = 0. No window opens until the next crossing (voltage lost). |
CMD.FLUSH | Close the open window (FREQ_OK = 0); none stays open. |
| Close while RES_VALID = 1 | Add into the shadow, max the peaks, MERGED = 1, FREQ_OK = 0. Energy is never lost. |
CMD.RES_ACK | Clear RES_VALID, FREQ_OK, FORCED, MERGED. |
| Accumulator add overflows | ACC_OVF = 1, sticky. |
Hand-over to the CPU
| # | Step | Detail |
|---|---|---|
| 1 | Copy | At window close the 7 accumulators (64 b each), N, peaks and the two zero-crossing latch sets copy to shadow registers. |
| 2 | Flag | RES_VALID = 1; level interrupt if MAC_IRQ_EN.WIN. The live accumulators keep running into the next window. |
| 3 | Read | CPU reads the shadow over APB (32-bit, 100 MHz clock enable, no async bridge). The shadow cannot change under it, so a 64-bit value cannot tear. |
| 4 | Acknowledge | CPU writes CMD.RES_ACK. If the next window closed first, its data was merged instead of lost. |
Use of the shadow registers, in two lines. The live accumulators change every 250 µs, and the CPU reads each 64-bit sum as two 32-bit halves (LO, then HI); if an add lands between the two reads, or an interrupt delays the CPU, it gets half-old, half-new garbage. The shadow is a frozen snapshot taken at the window edge, so the MAC keeps running and the CPU reads a stable result for as long as it needs.
Register map (base 0x4000_0000, APB, 32-bit)
| Offset | Name | Access | Fields |
|---|---|---|---|
| 0x000 | MAC_CTRL | RW | [0] EN, [1] AUTO |
| 0x004 | MAC_STATUS | RO | [0] RES_VALID, [1] FREQ_OK, [2] FORCED, [3] MERGED, [4] BLK_OVR, [5] ACC_OVF, [8] BUSY |
| 0x008 | MAC_CMD | WO | [0] START0, [1] START1, [2] FLUSH, [3] RES_ACK, [4] CLR |
| 0x00C | MAC_IRQ_EN | RW | [0] WIN, [1] ERR (BLK_OVR, ACC_OVF) |
| 0x010 – 0x018 | MAC_SRC_V0 / I0 / T0 | RW | SRAM address of half 0 per channel |
| 0x01C – 0x024 | MAC_SRC_V1 / I1 / T1 | RW | SRAM address of half 1 per channel |
| 0x028 | MAC_BLK_LEN | RW | samples per block per channel (80) |
| 0x02C | MAC_DLY | RW | [4:0] D_INT, [31:16] D_FRAC (Q16) |
| 0x030 | MAC_ZC_HYST | RW | [23:0] hysteresis |
| 0x034 | MAC_WIN_CFG | RW | [7:0] cycles, [23:8] max samples, [31:24] min gap |
| 0x040 – 0x074 | MAC_ACC_LO/HI(k) | RO | k = 0 … 6, shadow copies |
| 0x080 | MAC_NSAMP | RO | samples in the closed window |
| 0x084 / 0x088 | MAC_VPK / MAC_IPK | RO | peak |V|; peak max(|I|, |T|) |
| 0x08C – 0x094 | MAC_ZC0_IDX / PREV / CUR | RO | opening crossing: index, sample before (< 0), sample after (≥ 0) |
| 0x098 – 0x0A0 | MAC_ZC1_IDX / PREV / CUR | RO | closing crossing |
| 0x0A4 | MAC_WIN_CNT | RO | windows closed |
What firmware computes per window
| Quantity | Formula |
|---|---|
| V rms, I rms, T rms | K · sqrt(Σx² / N), K folds 1/1006633, PGA gain and the sensor ratio |
| Active power | Kv·Ki · Σv·i / N |
| Reactive power | Kv·Ki · Σv90·i / N |
| Apparent power, PF | S = Vrms · Irms, PF = P / S |
| Frequency | ≈ 10 × 4000 / N Hz, refined with the sub-sample fractions from ZC0 and ZC1. N = 800 → 50.000 Hz, N = 889 → 44.99 Hz. |
| Next delay | D = 1000 / f: 50 Hz → 20.000 (D_INT 20, D_FRAC 0); 49.5 Hz → 20.202 (D_INT 20, D_FRAC ≈ 0x33B8) |
| Energy, tariff, events | E += P · Δt; CF pulse per energy quantum; tamper, sag/swell, reverse-power and no-load checks |
| Measured (qemu, 1 s) | Software MAC | Hardware MAC |
|---|---|---|
| CPU instructions per second | 576,494 | 26,776 (−95.4 %) |
| CPU load at 200 MHz | 0.377 % | 0.022 % |
| Bill over 10 scenarios | reference | bit-exact |
The hardware-MAC figure works out to about 26,776 ÷ 5 ≈ 5,355 CPU instructions per window (derived here), including collecting the results and finishing the per-window arithmetic.
Reference
Timing and bandwidth in one place
| Event | Value | Note |
|---|---|---|
| Sample period | 250 µs | 4,000 sps |
| One 24-bit word on the wire | 11.72 µs | SCLK not in the Dolphin spec (O2); 2.048 MHz is an example |
| Synchroniser + word-done | < 1 µs | 2 flops at 100 MHz = 20 ns, plus one 24-clock word |
| FIFO slack per channel | 1 ms | 4 words assumed |
| DMA service budget | ≤ 15 µs | assumed |
| DMA half | 20 ms | 80 samples |
| Window | ≈ 200 ms | 10 cycles, 800 samples at 50 Hz; at most 889 samples = 222 ms |
| Results interrupts | 5 per second | one per window |
| CPU per window | ≈ 5,355 instr | derived from 26,776 per second |
| Flow | Rate | Basis |
|---|---|---|
| Serial, all 3 buses | 288 kbit/s | 72 bits × 4,000 |
| FIFO pushes | 12,000 words/s | 3 × 4,000 |
| DMA write to SRAM | 48 kB/s | 3 × 4 B × 4,000 |
| MAC read from SRAM | 48 kB/s | 3 × 4 B × 4,000 |
| SRAM total | 96 kB/s | 0.024 % of a 32-bit 100 MHz bus (400 MB/s) |
| CPU load | 0.022 % | at 200 MHz, hardware MAC build |
Reference
Where every number comes from
Each figure in this guide is a short calculation from a few starting facts. The starting facts are: MCLK 4.096 MHz, 4,000 sps, 50 Hz mains, 100 MHz MAC clock, 3 channels, 32-bit words. Two words are easy to mix up, so they are separated first.
| Word | Means | Length of one |
|---|---|---|
| Mains cycle | One full wave of the 50 Hz supply | 20 ms |
The other numbers
| Number | Steps | Result |
|---|---|---|
| Sampling | ||
| Modulator rate | MCLK 4.096 MHz ÷ 4 | 1.024 MHz |
| Sample rate | 1.024 MHz ÷ 256 (decimation) | 4,000 sps |
| Sample period | 1 ÷ 4,000 | 250 µs |
| Clocks per sample | 100 MHz ÷ 4,000 | 25,000 |
| Serial link and storage | ||
| Time for one word on the wire | 24 bits ÷ SCLK (example 2.048 MHz; real value from Dolphin, O2) | 11.7 µs |
| SCLK busy | 11.7 µs ÷ 250 µs | 4.7 % |
| Rate on the wire, per channel | 24 bits × 4,000 | 96 kb/s |
| Rate stored, per channel | 4 bytes × 4,000 | 16 kB/s |
| Rate into SRAM, all channels | 3 × 16 kB/s | 48 kB/s |
| SRAM port traffic | DMA write 48 + MAC read 48 | 96 kB/s |
| Interconnect traffic | 96 + DMA reading the FIFO 48 | 144 kB/s |
| Peak bus bandwidth | 4 bytes × 100 MHz | 400 MB/s |
| SRAM port use of the bus | 96 kB/s ÷ 400 MB/s | 0.024 % |
| FIFO slack | 4 words × 250 µs | 1 ms |
| Buffer | ||
| Samples per half | 4,000 ÷ 50 = samples in one mains cycle | 80 |
| Time to fill a half | 80 × 250 µs | 20 ms |
| Bytes per half | 80 samples × 3 channels × 4 B | 960 B |
| Both halves | 2 × 960 | 1,920 B |
| MAC work | ||
| Multiplies per sample | 7 products for the sums + 1 for the 90° delay | 8 |
| Multiplies per second | 8 × 4,000 | 32,000 |
| Window | ||
| Window time | 10 mains cycles × 20 ms | 200 ms |
| Samples per window | 200 ms ÷ 250 µs | 800 |
| Windows per second | 1 s ÷ 200 ms | 5 |
| Largest window | lowest allowed line 45 Hz: 10 × 4,000 ÷ 45 + 1 | 889 samples |
| Minimum crossing gap | highest allowed line 70 Hz: 4,000 ÷ 70 | 57 samples |
| Delay line depth | quarter wave at 45 Hz: 4,000 ÷ (4 × 45) = 22.2, + 2 taps, next power of 2 | 32 |
| Sum width | ||
| Largest input | 24-bit signed, |code| up to 2²³ | 2²³ |
| Largest product | 2²³ × 2²³ | 2⁴⁶ |
| Largest sum | 889 × 2⁴⁶ ≈ 2¹⁰ × 2⁴⁶ = 2⁵⁶, plus 1 sign bit | 57 bits |
| Headroom in 64 bits | 2⁶³ ÷ (889 × 2⁴⁶) = 131,072 ÷ 889 | 147 × |
| CPU | ||
| Instructions per window | 26,776 instructions per second ÷ 5 windows (derived from the measured figure) | ≈ 5,355 |
Reference
Reality check: what real meters do
To see which parts of this design are common practice and which are this project's own choices, they were compared with published documents from TI and Analog Devices and with the IEC 61000-4-30 power-quality standard. The verdict column says only what the sources checked actually show.
| Item | This design | What the sources show | Verdict |
|---|---|---|---|
| Active power | Σv·i, averaged over the window | TI MSP430 firmware: P = Σ v(n)·i(n) × scale, per frame. ADE9153A: multiplies the voltage and current waveforms, then low-pass filters. | matches |
| Reactive power | Σv90·i, with v90 made from a delay line plus a fractional step | TI: v90(n) is the voltage shifted by 90°, made of an integer delay of N samples plus a fractional delay filter. ADE9153A does the same idea the other way round: it shifts the current by 90° with a filter, then multiplies by the voltage. | matches TI method |
| RMS | Σv², Σi², Σt² averaged, then square root (firmware) | ADE9153A: square the signal, low-pass filter, take the square root. | matches |
| Sample rate | 4,000 samples per second | ADE9153A energy accumulators update at 4 kSPS. | same rate |
| Result registers | Shadow copy taken at window close | ADE9153A latches its 42-bit power accumulators into user registers at a programmable interval, from 500 µs to 2.048 s, or after a set number of half line cycles. | same idea |
| Accumulator width | 64 bits, plain sums over up to 889 samples (57 bits needed) | ADE9153A uses 42-bit accumulators fed by low-pass-filtered products, so it keeps running averages instead of raw window sums. | different, both valid |
| Window length | 10 mains cycles (200 ms at 50 Hz), started on a voltage zero-crossing | IEC 61000-4-30 Class A uses a basic interval of 10 cycles at 50 Hz (about 200 ms) and 12 cycles at 60 Hz. The standard joins these into gap-free intervals and resynchronises to the clock every 10 minutes. Starting each window on a zero-crossing is not required by it. | length matches; zero-crossing start is a design choice |
| Third channel T with Σt², Σv·t, Σv90·t | Neutral current summed like the phase current | Analog Devices ADE7953 (single voltage, two current channels) computes active and reactive power separately on both current channels (AWATT/BWATT, AVAR/BVAR). So V·I2 and V90·I2 per current channel is done in real meters. Not confirmed: Σt² as a separate sum, the neutral-current mismatch check, and doing all of it on one shared multiplier from SRAM (this project's own choices). | partly matches |
| Dedicated MAC that reads a buffered SRAM as an AXI master, filled by DMA ping-pong | Hardware does the per-sample sums, CPU does the per-window maths | Not found in the sources read. ADE9153A is a fixed DSP fed straight from the converter. TI does the sums in CPU firmware using its hardware multiplier. The Teridian 71M6541 has a separate compute engine, but its datasheet page read here gave no detail. | project choice, not verified against a product |
| Delay line of 32, 8 multiplies per sample, 80-sample halves | Internal sizing of this design | Derived on this page from 4,000 sps, 45 to 65 Hz and the 100 MHz clock. No outside source applies. | design numbers |
What this means for the manager conversation
The maths (Σv·i, a 90° shifted voltage, squares for RMS, latched results) is standard practice. The architecture (a bus-master MAC on buffered SRAM, and the split between hardware per sample and firmware per window) is this project's own decision, recorded in the decision log. Say it that way and do not claim other vendors do the same.
Sources
- TI SLAA494B, MSP430AFE2xx energy metering firmware calculations
- TI SLAA517F, MSP430F6736 watt-hour meter firmware
- Analog Devices ADE9153A technical reference manual (UG-1247)
- Summary of IEC 61000-4-30 Class A measurement intervals
The TI pages read did not include the RMS and apparent-power sections, and the Teridian datasheet excerpt had no compute-engine detail, so those two points are left unverified rather than assumed.
Reference
Decision log: each choice, the alternative, the reason
| Decision | Alternative | Reason |
|---|---|---|
| Dedicated capture slice | Stock SPI or I²S slave | SSYNC is one SCLK wide, data starts one SCLK after it, three buses run in parallel. A stock slave cannot frame that. |
| Sign-extend to 32 bits | 24-bit packing, zero-pad | Signed 32×32 multiplier and int32_t C arrays take the word unchanged; packing needs unaligned reads. |
| FIFO per channel | One interleaved FIFO | A dropped word in a shared queue swaps channel roles for every later sample; per-channel queues cannot. |
| Synchronous 100 MHz capture | Asynchronous FIFO per bus | SCLK ≤ MCLK 4.096 MHz gives ≥ 24× oversampling; one domain, 9 synchroniser cells to verify. |
| DMA | CPU interrupt per word | About 50–80 cycles per sample, 12,000 words/s: 2.5 % of a 32 MHz core at 4 kSps and 20 % at 32 kSps, no deep sleep, interrupt jitter smears sample timing. |
| SRAM ping-pong | FIFO straight into the MAC | Windows need ~800 samples resident together; the MAC clock can gate; the 90° delay needs history; samples stay for waveform and AI; the CPU sleeps about 19 of every 20 ms; no read/write collision. |
| Planar layout | Interleaved V, I, T | Three DMA channels each need one linear destination; the MAC has one base register per channel. |
| MAC as AXI master | AXI-slave MAC fed by the CPU | A CPU-fed MAC needs the CPU to move every operand across the bus; the master costs the CPU nothing per sample. |
| One shared multiplier | Seven parallel multipliers | 8 products in 25,000 clocks available; about 1/7 the area. |
| 64-bit accumulators + shadow | Read the live accumulators | 57 bits needed, 147× headroom; shadow copy gives torn-free reads and keeps the next window running. |
| Hardware 90° delay line | v90 buffer filled by firmware | Keeps all per-sample work in hardware; bit-exact against the 256-entry C reference; firmware only retunes the delay each window. |
| Zero-cross windows | Fixed 80-sample blocks | Whole mains cycles cancel the 2f ripple; fixed blocks would not when f ≠ 50.000 Hz. |
| Hardware/software split | All software, or all hardware | Per-sample work is fixed and regular (hardware). Per-window work runs 5 times a second and includes sqrt and division (firmware, 0.022 % of the CPU). |
| Clock-gated MAC | Always-clocked | The MAC needs only a few multiplies per 25,000 available clocks, so it is idle almost all the time. |
Reference
Open items
Open items found while writing this guide.
| ID | Item | Why it matters | Owner |
|---|---|---|---|
| O1 | MFE serial bit order: §5.4.2 text LSB-first, Fig 5.4 MSB-first | Wrong order gives a bit-reversed word with no error flag (try the simulator above). | Dolphin |
| O2 | Exact SCLK frequency. Not in the Dolphin spec; provable range is 100 kHz to 4.096 MHz | Sets synchroniser margin and word time. | Dolphin / DPHM |
| O3 | 200 MHz and 100 MHz from one PLL? | Decides the no-bridge, synchronous APB design. | SoC clock spec |
| O4 | Real DMA, capture and interrupt-controller register maps; SoC memory map | Firmware uses assumed maps; the SRAM base 0x2000_0000 here is assumed. | SoC RTL |
| O5 | Energy and CF pulse in hardware later? | Would move work out of the per-window firmware. | Architecture review |
| O6 | Half-done join: the MAC must wait until the V, I and T DMA channels have all finished | Otherwise the MAC can read a word the DMA has not written yet. | SoC RTL |
| O7 | Buffer in tightly-coupled memory, or SRAM read over AXI? | Sets the port and the read latency. | SoC RTL |
| O8 | FIFO overflow action and non-OKAY RRESP handling | Not defined; both should end in a sticky flag and the error interrupt. | MAC spec |
| O9 | AXI read latency (2 clocks) and ARID use are assumed | Confirm against the real interconnect and SRAM. | SoC RTL |