AFE-to-MAC Datapath
AFE-to-MAC datapath

Single-phase smart-meter SoC · study guide · 2026-09-29

AFE-to-MAC datapath

One sample of voltage, current and neutral current travels from the Dolphin AFE to the CPU in seven stages. Each stage below starts with a plain-words explanation, then gives the exact widths, addresses, handshakes and reasons. Every number is worked out on the page.

Source: Dolphin Metro-PM-MFE ViC specification R1.2

4,000samples per second, per channel
24 → 32bits per sample, before and after sign extension
80samples per buffer half (20 ms)
200 msone window: 10 mains waves, 800 samples

Basis of this guide

What this guide is based on

Tags: spec comes from the Dolphin specification, assumed is our own working choice, not yet confirmed, and open is unresolved (see the last section).

ParameterValueStatus
ScopeSingle phaseassumed
ChannelsV line voltage · I phase current · T neutral (tamper) currentassumed
AFEDolphin Metro-PM-MFE, serial DSP mode, 24-bit two's complementspec
Sample rate4,000 sps per channel, fixed. MCLK 4.096 MHz ÷ 4 = 1.024 MHz sync rate ÷ 256 = 4 kSpsspec
Stored word32-bit, sign-extended from 24-bitassumed
PathAFE → capture → FIFO → DMA → SRAM ping-pong → MAC reads SRAM as AXI masterassumed
MAC clock100 MHz, clock-gated, one shared 32×32→64 multiplierassumed
CPU clockUp to 200 MHz; 100 MHz = PLL 200 MHz ÷ 2, synchronousopen O3
Buffer2 halves × 3 channels × 80 samples × 4 B = 1,920 Bassumed
HW / SW splitMAC hardware: every per-sample operation. Firmware: per-window operations, 5 per secondassumed

Overview

The path in seven stages

The three channels stay in separate lanes (colour = channel) until they meet in the MAC. Pick a stage, or press Play. Every arrow shows its width and its rate. Change the Variables below the diagram and the labels, the tables and the numbers in every stage update.

Dolphin MFE V 24 b I 24 b T 24 b Capture V 24→32 I 24→32 T 24→32 FIFO V 32b ×4 I 32b ×4 T 32b ×4 DMA V ch0 I ch1 T ch2 SRAM, 2 halves T H0 T H1 I H0 I H1 V H0 V H1 MAC AXI rd, ZC 32×32→64 acc 7×64 Shadow regs acc shadow ZC latches RES_VALID RISC-V CPU APB read sqrt+scale billing ● V voltage ● I phase current ● T neutral current
Data lanes V, I, T from the AFE to the MAC. Between the AFE and capture each channel is three 1-bit wires (SCLK, SSYNC, SDATA), drawn as three thin lines. After capture every link is a 32-bit word. The DMA fills one SRAM half while the MAC reads the other. Lane order reverses in the SRAM column only so the three write routes do not cross.

What one word looks like at this stage (bit widths)

Variables

Assumed values are preset. Change any field to see a what-if; fields that differ from the assumed values are tagged. Nothing here changes the spec.

Results with the assumed values are: 48 kB/s into SRAM, 96 kB/s at the SRAM port (0.024 % of the bus), 57-bit sums, 147× headroom. The 25 result words per window are counted from the register map (my count). The MAC schedule is defined for 3 channels only, so the 7-channel option changes traffic and buffer sizes but reuses the same MAC cycles.

Stage 1 of 7 · AFE

AFE output: three serial words per sample instant

In plain words

The AFE (analog front end) is a separate Dolphin chip. It measures the voltage (V), the phase current (I) and the neutral current (T), and turns each into a 24-bit number. Every 250 µs there is a new number for each of the three.

It sends each number one bit at a time. Each channel uses 3 wires that carry one bit each: a bit clock (SCLK), a "new word starts" pulse (SSYNC) and the data (SDATA). Three channels make 9 wires. One bit at a time keeps the wire count small.

How fast SCLK runs is decided by the AFE and the Dolphin spec gives no number. It has to be at least 100 kHz (25 bits must fit in 250 µs) and cannot exceed the AFE's own 4.096 MHz clock.

Dolphin MFEexternal AFE chip3 sigma-delta ADCsdecimation ÷ 2564,000 sps, 24-bit wordslast 3 bits always 0MCLK 4.096 MHzANALOG INV input analogI input analogT input analogCONTROL (8-BIT BUS)MC_ADDR 8 bMC_DIN 8 bPER CHANNEL X = V, I, TMFE_SCLK_x 1 bMFE_SSYNC_x 1 bMFE_SDATA_x 1 bINTERRUPTSIRQ_V/I/T 3 × 1 b
The AFE as the rest of the design sees it. Pin names come from the Dolphin spec; the analog input names are not needed here.

"72 bits" is 3 channels × 24 bits. No 72-bit bus exists in the SoC: each channel has its own 3-wire serial bus, and each bus carries one 24-bit word per sample instant.

What leaves the AFE

ItemValueSource
ChannelsV unity-gain buffer, DO = 1006633 · Vin · 2 · I PGA ×4/8/16/32 · T PGA ×4/8/16/32Dolphin §1.1, §3.6
Modulator / sync rate1.024 MHz = MCLK 4.096 MHz ÷ 4§5.3 p.57
Decimation÷256 → 4,000 sps (MCLK ÷ 1024 overall)§3.4
Word24-bit two's complement, range −8,388,608 … +8,388,607 (0x800000 … 0x7FFFFF)§3.6
Real resolutionBits [2:0] are tied to 0 → code step 8, about 21 effective bits§5.4.2 p.59
Code → voltscurrent: Vin = code / (1006633 × gain)§3
Sampling instantSimultaneous: all three channels share one MCLK-synchronous sample edge§4.2.2
InterfaceDSP mode: MFE_SCLK, MFE_SSYNC, MFE_SDATA per channel (×3 buses), plus IRQ_V/I/T§2.2, §5.4
Bits per instant3 × 24 = 72 bits → 288 kbit/s at 4,000 sps
SCLK rateSet by the Dolphin AFE. The spec gives no value (§5.4 says only that the signals are synchronous to MCLK, 4.096 MHz). What the design can prove without it: SCLK must be at least 25 bits ÷ 250 µs = 100 kHz, and at most MCLK = 4.096 MHz. One 24-bit word then takes between 5.9 µs and 240 µs. The page uses 2.048 MHz (11.7 µs) only as an example valueopen O2
Bit order§5.4.2 text says LSB-first, Fig 5.4 shows MSB-first. The capture RTL must make this a parameter until Dolphin answers.open O1

Frame timing on one bus

MFE_SCLKMFE_SSYNCMFE_SDATA_x SSYNC: one SCLK wide 1 2 3 4 5 22 23 24 ··· Bits are numbered in transmit order. MFE drives on the falling edge;the receiver samples on each rising edge (dashed).
One 24-bit word on one bus. SSYNC is a single SCLK-wide pulse and data starts exactly one SCLK later. This is why a stock SPI or I²S peripheral cannot frame it.

Inside the AFE, before the pins

BlockJobParameter
Sensor + anti-alias RCShunt, CT or divider drives a differential pairRAA 1 kΩ, CAA 10 nF
PGAScales the sensor level to ADC full scaleI, T: ×4 … ×32 · V: fixed
ΔΣ modulator1-bit stream at 1.024 MHz carrying the signal in its densityorder and OSR not stated
Decimation filterRemoves out-of-band noise, drops the rate to Fs÷256 at 4 kSps
HPFRemoves DC and offsetcorner not specified
CalibrationGain and offset trim per channelGAIN_ERR_x, OFFSET_x
Phase shifterAligns V against I after sensor and filter delaysPSH_x: up to +17.9° at 50 Hz, step 0.0176°
SerialiserShifts the finished 24-bit word outsee frame timing above

Why this matters downstream: V and I are sampled at the same instant and phase-aligned inside the AFE. The SoC never re-aligns them; it only has to keep them index-aligned (sample n of V pairs with sample n of I). Every buffer choice below protects that.

RememberThe AFE hands over 24 bits, one at a time, 4,000 times a second, for each of V, I and T.

Stage 2 of 7 · Capture

Capture: three wires become three 32-bit words

In plain words

The capture block waits for the "new word" pulse, then shifts each arriving bit into a 24-bit register. After 24 clock ticks the number is complete. It then copies the top bit (the sign bit) into 8 new bits on the left, which makes a 32-bit number.

Why widen to 32 bits: the bus, the memory and the MAC all work in 32-bit words. Adding zeros on the left would turn a negative number into a huge positive one. Copying the sign bit keeps the value the same. This is called sign extension. Try it:

Try -5, then 5, then 8388607 (the biggest 24-bit value). The 8 grey bits on the left are always copies of the red sign bit.

Capture3 identical slices2-flop synchronisersSCLK rising-edge detectSSYNC arms, counter 0 to 2424-bit shift registersign-extend bit 23 to 32 bclk 100 MHzrst_nFROM AFE, PER SLICEsclk_x 1 b asyncssync_x 1 b asyncsdata_x 1 b asyncSETTINGlsb_first 1 b (O1)TO FIFO_Xpush_x 1 bword_x 32 b
Capture block. Three identical slices, one per channel. Nothing 72 bits wide exists. Port names are proposed: only the functions are known, not RTL names (open item O4).

A dedicated capture slice per channel turns the wiggling pins into one sign-extended 32-bit word. The three slices run side by side; their words complete within the same few clocks.

Mechanism, in order

#StepDetail
1SynchroniseSCLK, SSYNC and SDATA of each bus pass through 2 flip-flops into the 100 MHz domain. 3 buses × 3 signals = 9 synchroniser cells. Whatever Dolphin picks, SCLK ≤ 4.096 MHz, so the 100 MHz clock oversamples it at least 24×.
2Find the SCLK edgesclk_rise = sclk_s & ~sclk_s_d, one 100 MHz clock wide.
3Arm on SSYNCSSYNC seen → bit counter = 0, shifter armed. Data starts one SCLK later.
4ShiftOn each sclk_rise while armed: shift sdata_s into the 24-bit register, counter + 1.
5Word doneCounter = 24 → word_ready pulse, d[23:0] latched.
6Sign-extendword32 = {{8{d[23]}}, d[23:0]}
7PushWrite word32 into that channel's FIFO. Three slices push independently.

Watch one 24-bit sample arrive

Press Play. On every clock tick the AFE sends one bit and the capture block drops it into a 24-slot register. After 24 ticks the number is complete, and the top bit is copied to make a 32-bit word.

Sample:

1 · Bits arriving on the wire (first bit on the left)

2 · The 24-slot register (slot numbers below)

3 · The 32-bit word that goes to the FIFO

Advanced: sample value in hex, and what if the bit order is wrong?

The Dolphin spec contradicts itself on bit order (open item O1). Set the two menus to different orders and press Play: the number comes out bit-reversed and nothing in the datapath raises an error.

Why this matters: a wrong bit order gives a bit-reversed number with no error flag anywhere in the datapath. That is why O1 must be closed with Dolphin before RTL freeze, and why the receive order should be a register bit until then.

Sign extension

24-bit codeSigned value32-bit word storedName
0x7FFFF8+8,388,6000x007FFFF8highest value the AFE can send (low 3 bits are always 0)
0x000008+80x00000008smallest positive step shown here
0xFFFFF8−80xFFFFFFF8small negative
0x800000−8,388,6080xFF800000lowest value (most negative)
ChoiceReason
32-bit word, not 24SRAM and AXI are 32-bit; a 24-bit packing would need 3-byte addressing and unaligned reads.
Sign-extend, not zero-padThe MAC multiplier is signed 32×32. The word is already a valid int32_t, so neither the MAC nor the C code masks or extends anything. Zero-padding would turn −8 into +16,777,208.
Cost8 extra bits per word = 25 % storage overhead = 48 kB/s instead of 36 kB/s. Negligible.

Stage 3 of 7 · FIFO

FIFO: one small queue per channel

In plain words

The 32-bit word goes into a FIFO (first in, first out). Each channel has its own FIFO holding 4 words, and words leave in the order they came in.

Why a FIFO: the capture block produces a word when the AFE says so, and the DMA collects it when the bus is free. Those moments do not line up, so the FIFO holds the word in between. Four words at 4,000 samples per second is 1 ms of slack.

Why one per channel: if V, I and T shared one FIFO and one word were ever lost, everything after it would shift by one place and V would land in I's slot. Separate FIFOs keep a loss inside one channel. Three FIFOs is our assumed choice; one shared FIFO is the alternative.

FIFOone per channel (V, I, T)4 words × 32 bfirst in, first outa full FIFO sets overflowand drops the new wordclk 100 MHzrst_nFROM CAPTUREpush 1 bwdata 32 bFROM DMAread 1 bTO DMArdata 32 bdma_req 1 b = not emptySTATUSoverflow 1 b sticky
One FIFO per channel. Port names are proposed. The DMA reads the FIFO through a fixed address, and the request line is the "not empty" signal.
ParameterValueStatus
Count3: FIFO_V, FIFO_I, FIFO_Tassumed
Geometry32 bits wide × 4 words deep, 16 B each. Width follows the 32-bit word. Depth 4 is a margin choice with no published benchmark: keep it a parameter (DEPTH) and prove it with the stall test.assumed
ClockingSynchronous. Write side and read side are both in the 100 MHz domain; the clock-domain crossing was done by the 9 synchronisers in capture.assumed
Arrival rate1 word per channel per 250 µs (4,000 words/s each, 12,000 words/s total)
Slack4 words = 1 ms per channel. The steady-state need is under 1 word.assumed
DMA requestAsserted while the FIFO is not empty (one request line per stream)assumed
FlagsPer FIFO: empty, sticky overflowassumed
Overflow actionNot defined by Dolphin. Proposal: drop the new word, set the sticky flag, raise the error interrupt.open O8

Stage 4 of 7 · DMA and SRAM

DMA writes the words into SRAM ping-pong halves

In plain words

The DMA (direct memory access) is a small hardware courier. It takes each word out of a FIFO and writes it into SRAM, then moves on 4 bytes to the next slot. The CPU does none of this.

SRAM is split into two halves of 80 samples each. The DMA writes into one half, and the MAC reads the other half, which is already full. When the DMA finishes its half, the two switch jobs. This is called ping-pong. Filling a half takes 80 ÷ 4,000 = 20 ms.

The MAC is fast. When a half is full, the DMA sends a half_done event. The MAC then reads all 80 samples of V, I and T and goes idle. The DMA needs about 20 ms to fill the other half, so the MAC always finishes long before that half is full.

What switches is the address. The DMA moves its write address to the other half by itself, so it never stops. The MAC is not moved: it is started on the half that just filled. half_done for half 0 means "read the half 0 addresses", then half 1, then half 0 again.

DMA3 metering channelssource fixed, dest + 4 B80 samples per halfping-pong swap at boundarypriority: metering firstclkrst_nREQUESTSreq_v/i/t 3 × 1 bSETUP OVER APBsrc, dst, length regslock 1 b per chAXI MASTER, 32 Bread FIFO reg AR / Rwrite SRAM AW / W / BEVENTShalf_done 1 b to MACoverrun 1 b per ch
The three metering channels of the DMA. Port names are proposed. The real register map is open item O4.
SRAMping-pong bufferhalf 0 and half 1each half: V, I, T × 80 words960 B per half, 1,920 B totaladdr = base + 0x3C0·h + 0x140·c + 4·nclk 100 MHzAXI SLAVE, 32 Bwrite from DMAread address from MACAXI SLAVE, 32 Bread data to MACresponse OKAY / error
The SRAM buffer as the DMA and MAC see it. The base address 0x2000_0000 is assumed.
TimeDMA writesMAC reads
0 to 20 mshalf 0idle
20 msswitches to half 1, sends half_done for half 0reads half 0
20 to 40 mshalf 1idle
40 msswitches back to half 0, sends half_done for half 1reads half 1

If the MAC were still busy when the next half_done arrived, the BLK_OVR flag would be set so the problem is visible. That cannot happen while the MAC has 20 ms and needs only a tiny fraction of it.

DMA channels

Fieldch0 (V)ch1 (I)ch2 (T)
SourceFIFO_V data reg, fixedFIFO_I data reg, fixedFIFO_T data reg, fixed
DestinationV region, +4 per wordI region, +4 per wordT region, +4 per word
Width32 bit32 bit32 bit
Count80 words per half8080
ModePing-pong, auto-swapsamesame
TriggerFIFO_V not emptyFIFO_I not emptyFIFO_T not empty
AspectValueStatus
Per trigger1 AXI read of the FIFO register and 1 AXI write to SRAM, 4 bytes each: ARLEN=0, ARSIZE=2, AWLEN=0, AWSIZE=2, WSTRB=4'b1111
Interrupt granularityOne event per 80 words, not per word
Latency budgetRequest to bus transaction ≤ 15 µs under full contention (specified at 32 kSps; at 4 kSps the sample period is 250 µs). A requirement, not yet shown to be met by the bus design.assumed
PriorityMetering A above comms B above memory copy C, pre-empted at beat boundaries
OverrunPer-channel flag if the write pointer laps unread data
ProtectionMetering channels software-lockable
DMA traffic3 × 4 B × 4,000 = 48 kB/s written to SRAM, 48 kB/s read from FIFOscomputed

Wait for all three channels. V, I and T each have their own DMA channel, and they finish their 80th word a few microseconds apart. Say V finishes first and T last. If the MAC started as soon as V was done, it would reach the end of the T data before the DMA had written the last T word, and it would read an old value. So the MAC must wait until all three channels have finished (V and I and T), and only then start. The design notes say "half-done starts the MAC" but do not say this wait is built in, so it is open item O6.

SRAM layout

Planar: each channel is one contiguous array of 80 words per half. That matches three independent DMA channels with one incrementing destination each, and one source-address register per channel in the MAC. A whole window of a single channel is also contiguous for later waveform or AI use. Base address 0x2000_0000 is assumed (the SoC memory map is not defined, O4).

RegionAddress rangeBytes (one channel, one half)MAC register
V half 00x2000_0000 – 0x2000_013F320MAC_SRC_V0
I half 00x2000_0140 – 0x2000_027F320MAC_SRC_I0
T half 00x2000_0280 – 0x2000_03BF320MAC_SRC_T0
V half 10x2000_03C0 – 0x2000_04FF320MAC_SRC_V1
I half 10x2000_0500 – 0x2000_063F320MAC_SRC_I1
T half 10x2000_0640 – 0x2000_077F320MAC_SRC_T1
Total, 6 rows1,920MAC_BLK_LEN = 80

One row above is one channel in one half: 80 samples × 4 B = 320 B. A channel has two rows (half 0 and half 1), so one channel uses 2 × 320 = 640 B in total. One half holds three channels: 3 × 320 = 960 B (this is the 0x3C0 step between halves). Both halves: 1,920 B.

Address of sample n: base + half × 0x3C0 + channel × 0x140 + n × 4, with channel 0 = V, 1 = I, 2 = T and n = 0 … 79. Memory is little-endian: a stored word holds d[7:0] at the lowest address, then d[15:8], d[23:16], and last the sign byte 0x00 or 0xFF.

Address calculator

Ping-pong timing

DMA writesMAC reads 0 ms 20 ms 40 ms 60 ms 80 ms fills half 0 fills half 1 fills half 0 fills half 1 half 0 donereads half 0 half 1 donereads half 1 half 0 donereads half 0 Bar widths: the DMA blocks are to scale. The MAC bars are exaggerated: the MAC needs only a tiny fraction of a 20 ms half. The MAC clock is gated (off) between bars.
The DMA fills one half for 20 ms (80 samples at 250 µs). At the boundary the halves swap and the MAC reads the half that just completed while the DMA fills the other.
EventEffect
Half done (all 3 channels)MAC starts on that half (MAC_CTRL.AUTO = 1). DMA continues into the other half.
Half done while MAC still busyMAC_STATUS.BLK_OVR = 1, sticky. The MAC needs a tiny fraction of the 20 ms half, so this only triggers on a fault.
Why two halvesThe MAC is never reading a slot the DMA is writing, without any lock or copy.

Stage 5 of 7 · MAC read side

The MAC reads SRAM as an AXI master

In plain words

The MAC does not wait to be handed data. It sends its own read request: an address goes out on the bus and SRAM answers with one 32-bit word. That makes the MAC the bus master.

For every sample it makes 3 reads (V, I, T). The data is read straight from the buffer where the DMA left it, so it is never copied a second time.

MAC read portAXI master3 reads per sampleaddress = SRC + 4 · nV, then I, then Twaits for RVALIDclk_mac (gated)FROM SRAMRDATA 32 bRVALID 1 bRRESP 2 bARREADY 1 bTO SRAMARADDR 32 bARVALID 1 bARLEN 0, single beatARSIZE 2, 4 bytes
The read side of the MAC. ARID is assumed (0 = V, 1 = I, 2 = T) and not shown. See open item O9.

The MAC generates its own read addresses. Nothing pushes operands to it: once started it walks the half sample by sample, three single-beat reads per sample (V, I, T).

Sequence per block

#StepDetail
1StartAll-channel half-done event with MAC_CTRL.AUTO = 1, or MAC_CMD.START0/START1 written by the CPU. Clock gate opens.
2Pick the halfThe half number selects MAC_SRC_V/I/T[h], three 32-bit base addresses.
3AddressSample counter n = 0 … MAC_BLK_LEN−1 (80). ARADDR = SRC_ch[h] + 4·n for each of V, I, T.
4ReadIssue AR_V(n), AR_I(n), AR_T(n) back to back; accept three R beats.
5ComputeThe three words are the inputs for the seven sums (next stage).
6PrefetchReads for sample n+1 are issued while sample n is still being processed, so the multiplier does not wait for the bus.
7Finishn reaches 80 → BUSY clears, clock gate closes.
clk 0 clk 1 clk 2 clk 3 clk 4 clk 5 clk 6 clk 7 ACLK ARVALID ARREADY ARADDR RVALID RDATA V[5] I[5] T[5] dV dI dT ARREADY = 1: slave always ready (assumed) 3 requests, 1 per clock data 2 clocks after accept MAC starts on this sample
Three single-beat reads for sample n = 5 of half 0: ARADDR = V[5] 0x2000_0014, I[5] 0x2000_0154, T[5] 0x2000_0294. A request is accepted on the clock where ARVALID and ARREADY are both high.

AXI signals used

ChannelSignalWidthValue from the MAC
ARARADDR32byte address, word-aligned (bits [1:0] = 00)
ARARLEN80 → one beat
ARARSIZE33'b010 → 4 bytes
ARARBURST22'b01 INCR (no effect at length 1)
ARARID2assumed 0 = V, 1 = I, 2 = T, so returns can be matched
ARARVALID / ARREADY1 / 1Master drives VALID and holds address until READY; VALID must not wait for READY
RRDATA32the stored sign-extended word
RRRESP2expected 2'b00 OKAY. Handling of any other value is not defined (O8).
RRLAST11 on the only beat
RRVALID / RREADY1 / 1MAC holds RREADY high whenever it expects data
ParameterValue
Read traffic3 × 4 B × 4,000 = 48 kB/s
SRAM traffic totalDMA write 48 kB/s + MAC read 48 kB/s = 96 kB/s = 0.024 % of a 32-bit 100 MHz bus (400 MB/s)
Read latency2 clocks after accept assumed (synchronous SRAM); hidden by prefetch
ClockMAC, DMA, capture and the bus in this path are all 100 MHz, synchronous. No clock-domain bridge on the sample path.
Memory portThe buffers could sit in tightly-coupled memory or in SRAM read over AXI. Which port and latency apply is open (O7).

Why the MAC is the master

Suppose the CPU fed the MAC instead. Each operand would be read from SRAM by the CPU and written to the MAC, so it would cross the interconnect twice, and the CPU would run a loop 4,000 times a second. With the MAC as master, each operand is read once, by the MAC itself. The CPU has nothing to do per sample and can sleep between windows.

RememberThe MAC pulls its own data from SRAM, 3 single-word reads per sample.

Stage 6 of 7 · MAC compute

Seven accumulations per sample

In plain words

The MAC (multiply and accumulate) is hardware that runs by itself. For every sample it multiplies and adds into 7 running totals: V×V, I×I, T×T, V×I, V×T, and two more that use a copy of the voltage delayed by a quarter wave (used for reactive power).

The CPU only sets it up (where the buffers are, how long) and switches it on. The multiplying is wired hardware, not program instructions. The totals are 64 bits wide because a product is up to 46 bits and up to 889 are added in one window, which needs 57 bits.

The MAC has 25,000 clocks per sample (100 MHz ÷ 4,000 sps) and needs only 8 multiplies, so it is idle almost all the time and its clock is gated off between blocks.

MACcompute enginesample registers V, I, Tdelay line 32 × 32 binterpolate to make v90zero-cross with hysteresis1 multiplier 32×32 to 647 accumulators × 64 bpeak detect, window framerclk_macrst_nSTART AND SETUPhalf_done from DMACTRL, CMD APBSRC_x, BLK_LEN APBDLY, ZC_HYST APBWIN_CFG APBSAMPLESV, I, T words 3 × 32 bREADSARADDR etc. to SRAMRESULTSwindow_close 1 bsums, peaks to shadowSTATUSBUSY flagBLK_OVR stickyACC_OVF sticky
The MAC block with its outside connections and the parts inside it. Register names come from the project register map.

Inputs and internal state

ItemWidthMeaning
v, i, t32 signedthe three words just read
Delay line32 × 32circular register file, one v written per sample
D_INT5MAC_DLY[4:0]: whole-sample part of the 90° delay
D_FRAC16MAC_DLY[31:16]: fractional part, Q16 (0x8000 = 0.5)
x0, x132 signedtwo adjacent delay-line taps around the delay
v9032 signedx0 + (((x1 − x0) × D_FRAC) >>> 16), voltage delayed by a quarter cycle
QuantityValue
Delay neededFs / (4 f) = 1000 / f samples: 20.000 at 50 Hz, 22.2 at 45 Hz (worst case, so the line is 32 deep)
Why 90°Reactive power is the average of v(t − T/4) × i(t). A quarter period is an exact integer only at 50.000 Hz, so the fraction is interpolated between two neighbouring samples.
Who sets the delayFirmware, from the measured frequency, once per window. No hardware divider.

The seven accumulators

kNameSumBecomesRegister (assumed order)
0Σv²v × vV rmsACC_LO/HI(0) 0x040 / 0x044
1Σi²i × iI rms0x048 / 0x04C
2Σt²t × tT (neutral) rms0x050 / 0x054
3Σv·iv × iactive power, phase0x058 / 0x05C
4Σv·tv × tactive power, neutral0x060 / 0x064
5Σv90·iv90 × ireactive power, phase0x068 / 0x06C
6Σv90·tv90 × treactive power, neutral0x070 / 0x074

Accumulator width

StepValue
Largest input|code| = 2²³ (at 0x800000)
Largest product2²³ × 2²³ = 2⁴⁶ → 48-bit signed
Largest window889 samples (10 cycles at 45 Hz)
Largest sum889 × 2⁴⁶ = 6.26 × 10¹⁶ ≈ 2⁵⁵·⁸ → 57 bits signed
Register64-bit signed → 2⁶³ / 6.26 × 10¹⁶ = 147× headroom (2¹⁷ = 131,072 samples, 32.8 s, before any overflow)
Why 64, not 57The multiplier produces 64 bits, and a 64-bit value reads cleanly as two 32-bit words over APB.
If it ever overflowsMAC_STATUS.ACC_OVF = 1, sticky

Cost of the per-sample work

QuantityCalculationValue
Clocks available per sample100 MHz ÷ 4,00025,000
Multiplies per sample7 sums + 1 for v908
Multipliers1 shared instead of 7 parallel≈ 1/7 area

8 multiplies against 25,000 available clocks: timing is not a design constraint. How many clocks the MAC actually needs depends on the multiplier and control logic chosen, and is not fixed here.

Stage 7 of 7 · Window close

Windows close on zero crossings, then the CPU takes over

In plain words

A window is the group of samples that makes one result: exactly 10 mains waves. At 50 Hz that is 10 × 20 ms = 200 ms, or 800 samples. It is not the same as a buffer half: a half is 80 samples of memory, so one window spans 10 halves.

At the end of a window the MAC copies its 7 totals into shadow registers, raises a flag and starts the next window straight away. The CPU reads a 64-bit total as two 32-bit halves, and an add in between would give half-old, half-new data. The frozen shadow copy avoids that.

The CPU then does the last steps from the shadow totals: square root for RMS, scaling to volts and amps, energy and events. That happens 5 times a second. The window is set to end on a voltage zero-crossing so it always covers whole waves.

Result registersshadow copy + APBfrozen copy at window closebase 0x4000_0000read as 2 × 32 b halvesRES_ACK clears RES_VALIDnext window merges if lateclk 100 MHzrst_nFROM MACwindow_close 1 b7 sums 7 × 64 bN, peaks, ZC regsAPB SLAVE, 32 BPSEL, PENABLE 1 b eachPWRITE, PADDR offsetPWDATA 32 bTO CPUPRDATA 32 birq_win 1 b levelSTATUSRES_VALID 1 bMERGED, FORCED 1 b each
The result registers the CPU reads. The APB address width is not defined; the offsets run from 0x000 to 0x0A4.

10 here means 10 mains cycles (10 waves of the 50 Hz supply = 200 ms), not 10 MAC clock ticks. See Numbers.

Windows are not the same thing as buffer halves. The DMA halves are fixed at 80 samples; a window is 10 mains cycles found by the MAC from the voltage zero crossings, about 800 samples, spanning roughly ten halves. Averaging over whole cycles cancels the 2f ripple in v·i and v² that a fixed block would leave when the line is not exactly 50.000 Hz.

Window configuration

FieldRegisterValueDerivation
Cycles per windowMAC_WIN_CFG[7:0]1010 cycles = 200 ms at 50 Hz
Max samplesMAC_WIN_CFG[23:8]88910 × 4,000 ÷ 45 + 1
Min crossing gapMAC_WIN_CFG[31:24]574,000 ÷ 70 Hz, rejects double crossings
Re-arm hysteresisMAC_ZC_HYST[23:0]codesa crossing re-arms only after |v| exceeds it

Behaviour rules

EventAction
Positive crossing, armed, ≥ min gap since the lastNo window open: open one and latch ZC0. Otherwise count it; on the 10th, close (FREQ_OK = 1, latch ZC1) and open the next (ZC0 = this crossing).
First crossing after start or after a forced closeClose whatever accumulated with FREQ_OK = 0, then open a window.
N reaches max samples (889)Close with FORCED = 1, FREQ_OK = 0. No window opens until the next crossing (voltage lost).
CMD.FLUSHClose the open window (FREQ_OK = 0); none stays open.
Close while RES_VALID = 1Add into the shadow, max the peaks, MERGED = 1, FREQ_OK = 0. Energy is never lost.
CMD.RES_ACKClear RES_VALID, FREQ_OK, FORCED, MERGED.
Accumulator add overflowsACC_OVF = 1, sticky.

Hand-over to the CPU

#StepDetail
1CopyAt window close the 7 accumulators (64 b each), N, peaks and the two zero-crossing latch sets copy to shadow registers.
2FlagRES_VALID = 1; level interrupt if MAC_IRQ_EN.WIN. The live accumulators keep running into the next window.
3ReadCPU reads the shadow over APB (32-bit, 100 MHz clock enable, no async bridge). The shadow cannot change under it, so a 64-bit value cannot tear.
4AcknowledgeCPU writes CMD.RES_ACK. If the next window closed first, its data was merged instead of lost.

Use of the shadow registers, in two lines. The live accumulators change every 250 µs, and the CPU reads each 64-bit sum as two 32-bit halves (LO, then HI); if an add lands between the two reads, or an interrupt delays the CPU, it gets half-old, half-new garbage. The shadow is a frozen snapshot taken at the window edge, so the MAC keeps running and the CPU reads a stable result for as long as it needs.

Register map (base 0x4000_0000, APB, 32-bit)

OffsetNameAccessFields
0x000MAC_CTRLRW[0] EN, [1] AUTO
0x004MAC_STATUSRO[0] RES_VALID, [1] FREQ_OK, [2] FORCED, [3] MERGED, [4] BLK_OVR, [5] ACC_OVF, [8] BUSY
0x008MAC_CMDWO[0] START0, [1] START1, [2] FLUSH, [3] RES_ACK, [4] CLR
0x00CMAC_IRQ_ENRW[0] WIN, [1] ERR (BLK_OVR, ACC_OVF)
0x010 – 0x018MAC_SRC_V0 / I0 / T0RWSRAM address of half 0 per channel
0x01C – 0x024MAC_SRC_V1 / I1 / T1RWSRAM address of half 1 per channel
0x028MAC_BLK_LENRWsamples per block per channel (80)
0x02CMAC_DLYRW[4:0] D_INT, [31:16] D_FRAC (Q16)
0x030MAC_ZC_HYSTRW[23:0] hysteresis
0x034MAC_WIN_CFGRW[7:0] cycles, [23:8] max samples, [31:24] min gap
0x040 – 0x074MAC_ACC_LO/HI(k)ROk = 0 … 6, shadow copies
0x080MAC_NSAMPROsamples in the closed window
0x084 / 0x088MAC_VPK / MAC_IPKROpeak |V|; peak max(|I|, |T|)
0x08C – 0x094MAC_ZC0_IDX / PREV / CURROopening crossing: index, sample before (< 0), sample after (≥ 0)
0x098 – 0x0A0MAC_ZC1_IDX / PREV / CURROclosing crossing
0x0A4MAC_WIN_CNTROwindows closed

What firmware computes per window

QuantityFormula
V rms, I rms, T rmsK · sqrt(Σx² / N), K folds 1/1006633, PGA gain and the sensor ratio
Active powerKv·Ki · Σv·i / N
Reactive powerKv·Ki · Σv90·i / N
Apparent power, PFS = Vrms · Irms, PF = P / S
Frequency≈ 10 × 4000 / N Hz, refined with the sub-sample fractions from ZC0 and ZC1. N = 800 → 50.000 Hz, N = 889 → 44.99 Hz.
Next delayD = 1000 / f: 50 Hz → 20.000 (D_INT 20, D_FRAC 0); 49.5 Hz → 20.202 (D_INT 20, D_FRAC ≈ 0x33B8)
Energy, tariff, eventsE += P · Δt; CF pulse per energy quantum; tamper, sag/swell, reverse-power and no-load checks
Measured (qemu, 1 s)Software MACHardware MAC
CPU instructions per second576,49426,776 (−95.4 %)
CPU load at 200 MHz0.377 %0.022 %
Bill over 10 scenariosreferencebit-exact

The hardware-MAC figure works out to about 26,776 ÷ 5 ≈ 5,355 CPU instructions per window (derived here), including collecting the results and finishing the per-window arithmetic.

RememberA window is 10 mains waves. It gives one set of totals, the shadow copy keeps them steady, and the CPU turns them into volts, amps and energy.

Reference

Timing and bandwidth in one place

EventValueNote
Sample period250 µs4,000 sps
One 24-bit word on the wire11.72 µsSCLK not in the Dolphin spec (O2); 2.048 MHz is an example
Synchroniser + word-done< 1 µs2 flops at 100 MHz = 20 ns, plus one 24-clock word
FIFO slack per channel1 ms4 words assumed
DMA service budget≤ 15 µsassumed
DMA half20 ms80 samples
Window≈ 200 ms10 cycles, 800 samples at 50 Hz; at most 889 samples = 222 ms
Results interrupts5 per secondone per window
CPU per window≈ 5,355 instrderived from 26,776 per second
FlowRateBasis
Serial, all 3 buses288 kbit/s72 bits × 4,000
FIFO pushes12,000 words/s3 × 4,000
DMA write to SRAM48 kB/s3 × 4 B × 4,000
MAC read from SRAM48 kB/s3 × 4 B × 4,000
SRAM total96 kB/s0.024 % of a 32-bit 100 MHz bus (400 MB/s)
CPU load0.022 %at 200 MHz, hardware MAC build

Reference

Where every number comes from

Each figure in this guide is a short calculation from a few starting facts. The starting facts are: MCLK 4.096 MHz, 4,000 sps, 50 Hz mains, 100 MHz MAC clock, 3 channels, 32-bit words. Two words are easy to mix up, so they are separated first.

WordMeansLength of one
Mains cycleOne full wave of the 50 Hz supply20 ms

The other numbers

NumberStepsResult
Sampling
Modulator rateMCLK 4.096 MHz ÷ 41.024 MHz
Sample rate1.024 MHz ÷ 256 (decimation)4,000 sps
Sample period1 ÷ 4,000250 µs
Clocks per sample100 MHz ÷ 4,00025,000
Serial link and storage
Time for one word on the wire24 bits ÷ SCLK (example 2.048 MHz; real value from Dolphin, O2)11.7 µs
SCLK busy11.7 µs ÷ 250 µs4.7 %
Rate on the wire, per channel24 bits × 4,00096 kb/s
Rate stored, per channel4 bytes × 4,00016 kB/s
Rate into SRAM, all channels3 × 16 kB/s48 kB/s
SRAM port trafficDMA write 48 + MAC read 4896 kB/s
Interconnect traffic96 + DMA reading the FIFO 48144 kB/s
Peak bus bandwidth4 bytes × 100 MHz400 MB/s
SRAM port use of the bus96 kB/s ÷ 400 MB/s0.024 %
FIFO slack4 words × 250 µs1 ms
Buffer
Samples per half4,000 ÷ 50 = samples in one mains cycle80
Time to fill a half80 × 250 µs20 ms
Bytes per half80 samples × 3 channels × 4 B960 B
Both halves2 × 9601,920 B
MAC work
Multiplies per sample7 products for the sums + 1 for the 90° delay8
Multiplies per second8 × 4,00032,000
Window
Window time10 mains cycles × 20 ms200 ms
Samples per window200 ms ÷ 250 µs800
Windows per second1 s ÷ 200 ms5
Largest windowlowest allowed line 45 Hz: 10 × 4,000 ÷ 45 + 1889 samples
Minimum crossing gaphighest allowed line 70 Hz: 4,000 ÷ 7057 samples
Delay line depthquarter wave at 45 Hz: 4,000 ÷ (4 × 45) = 22.2, + 2 taps, next power of 232
Sum width
Largest input24-bit signed, |code| up to 2²³2²³
Largest product2²³ × 2²³2⁴⁶
Largest sum889 × 2⁴⁶ ≈ 2¹⁰ × 2⁴⁶ = 2⁵⁶, plus 1 sign bit57 bits
Headroom in 64 bits2⁶³ ÷ (889 × 2⁴⁶) = 131,072 ÷ 889147 ×
CPU
Instructions per window26,776 instructions per second ÷ 5 windows (derived from the measured figure)≈ 5,355

Reference

Reality check: what real meters do

To see which parts of this design are common practice and which are this project's own choices, they were compared with published documents from TI and Analog Devices and with the IEC 61000-4-30 power-quality standard. The verdict column says only what the sources checked actually show.

ItemThis designWhat the sources showVerdict
Active powerΣv·i, averaged over the windowTI MSP430 firmware: P = Σ v(n)·i(n) × scale, per frame. ADE9153A: multiplies the voltage and current waveforms, then low-pass filters.matches
Reactive powerΣv90·i, with v90 made from a delay line plus a fractional stepTI: v90(n) is the voltage shifted by 90°, made of an integer delay of N samples plus a fractional delay filter. ADE9153A does the same idea the other way round: it shifts the current by 90° with a filter, then multiplies by the voltage.matches TI method
RMSΣv², Σi², Σt² averaged, then square root (firmware)ADE9153A: square the signal, low-pass filter, take the square root.matches
Sample rate4,000 samples per secondADE9153A energy accumulators update at 4 kSPS.same rate
Result registersShadow copy taken at window closeADE9153A latches its 42-bit power accumulators into user registers at a programmable interval, from 500 µs to 2.048 s, or after a set number of half line cycles.same idea
Accumulator width64 bits, plain sums over up to 889 samples (57 bits needed)ADE9153A uses 42-bit accumulators fed by low-pass-filtered products, so it keeps running averages instead of raw window sums.different, both valid
Window length10 mains cycles (200 ms at 50 Hz), started on a voltage zero-crossingIEC 61000-4-30 Class A uses a basic interval of 10 cycles at 50 Hz (about 200 ms) and 12 cycles at 60 Hz. The standard joins these into gap-free intervals and resynchronises to the clock every 10 minutes. Starting each window on a zero-crossing is not required by it.length matches; zero-crossing start is a design choice
Third channel T with Σt², Σv·t, Σv90·tNeutral current summed like the phase currentAnalog Devices ADE7953 (single voltage, two current channels) computes active and reactive power separately on both current channels (AWATT/BWATT, AVAR/BVAR). So V·I2 and V90·I2 per current channel is done in real meters. Not confirmed: Σt² as a separate sum, the neutral-current mismatch check, and doing all of it on one shared multiplier from SRAM (this project's own choices).partly matches
Dedicated MAC that reads a buffered SRAM as an AXI master, filled by DMA ping-pongHardware does the per-sample sums, CPU does the per-window mathsNot found in the sources read. ADE9153A is a fixed DSP fed straight from the converter. TI does the sums in CPU firmware using its hardware multiplier. The Teridian 71M6541 has a separate compute engine, but its datasheet page read here gave no detail.project choice, not verified against a product
Delay line of 32, 8 multiplies per sample, 80-sample halvesInternal sizing of this designDerived on this page from 4,000 sps, 45 to 65 Hz and the 100 MHz clock. No outside source applies.design numbers

What this means for the manager conversation

The maths (Σv·i, a 90° shifted voltage, squares for RMS, latched results) is standard practice. The architecture (a bus-master MAC on buffered SRAM, and the split between hardware per sample and firmware per window) is this project's own decision, recorded in the decision log. Say it that way and do not claim other vendors do the same.

Sources

The TI pages read did not include the RMS and apparent-power sections, and the Teridian datasheet excerpt had no compute-engine detail, so those two points are left unverified rather than assumed.

Reference

Decision log: each choice, the alternative, the reason

DecisionAlternativeReason
Dedicated capture sliceStock SPI or I²S slaveSSYNC is one SCLK wide, data starts one SCLK after it, three buses run in parallel. A stock slave cannot frame that.
Sign-extend to 32 bits24-bit packing, zero-padSigned 32×32 multiplier and int32_t C arrays take the word unchanged; packing needs unaligned reads.
FIFO per channelOne interleaved FIFOA dropped word in a shared queue swaps channel roles for every later sample; per-channel queues cannot.
Synchronous 100 MHz captureAsynchronous FIFO per busSCLK ≤ MCLK 4.096 MHz gives ≥ 24× oversampling; one domain, 9 synchroniser cells to verify.
DMACPU interrupt per wordAbout 50–80 cycles per sample, 12,000 words/s: 2.5 % of a 32 MHz core at 4 kSps and 20 % at 32 kSps, no deep sleep, interrupt jitter smears sample timing.
SRAM ping-pongFIFO straight into the MACWindows need ~800 samples resident together; the MAC clock can gate; the 90° delay needs history; samples stay for waveform and AI; the CPU sleeps about 19 of every 20 ms; no read/write collision.
Planar layoutInterleaved V, I, TThree DMA channels each need one linear destination; the MAC has one base register per channel.
MAC as AXI masterAXI-slave MAC fed by the CPUA CPU-fed MAC needs the CPU to move every operand across the bus; the master costs the CPU nothing per sample.
One shared multiplierSeven parallel multipliers8 products in 25,000 clocks available; about 1/7 the area.
64-bit accumulators + shadowRead the live accumulators57 bits needed, 147× headroom; shadow copy gives torn-free reads and keeps the next window running.
Hardware 90° delay linev90 buffer filled by firmwareKeeps all per-sample work in hardware; bit-exact against the 256-entry C reference; firmware only retunes the delay each window.
Zero-cross windowsFixed 80-sample blocksWhole mains cycles cancel the 2f ripple; fixed blocks would not when f ≠ 50.000 Hz.
Hardware/software splitAll software, or all hardwarePer-sample work is fixed and regular (hardware). Per-window work runs 5 times a second and includes sqrt and division (firmware, 0.022 % of the CPU).
Clock-gated MACAlways-clockedThe MAC needs only a few multiplies per 25,000 available clocks, so it is idle almost all the time.

Reference

Open items

Open items found while writing this guide.

IDItemWhy it mattersOwner
O1MFE serial bit order: §5.4.2 text LSB-first, Fig 5.4 MSB-firstWrong order gives a bit-reversed word with no error flag (try the simulator above).Dolphin
O2Exact SCLK frequency. Not in the Dolphin spec; provable range is 100 kHz to 4.096 MHzSets synchroniser margin and word time.Dolphin / DPHM
O3200 MHz and 100 MHz from one PLL?Decides the no-bridge, synchronous APB design.SoC clock spec
O4Real DMA, capture and interrupt-controller register maps; SoC memory mapFirmware uses assumed maps; the SRAM base 0x2000_0000 here is assumed.SoC RTL
O5Energy and CF pulse in hardware later?Would move work out of the per-window firmware.Architecture review
O6Half-done join: the MAC must wait until the V, I and T DMA channels have all finishedOtherwise the MAC can read a word the DMA has not written yet.SoC RTL
O7Buffer in tightly-coupled memory, or SRAM read over AXI?Sets the port and the read latency.SoC RTL
O8FIFO overflow action and non-OKAY RRESP handlingNot defined; both should end in a sticky flag and the error interrupt.MAC spec
O9AXI read latency (2 clocks) and ARID use are assumedConfirm against the real interconnect and SRAM.SoC RTL