case 13 · Personal · 2026
FPGA Human Presence Detector — Verilog CNN, binary classifier
A four-layer convolutional neural network built entirely in Verilog on a Tang Primer 20K — no CPU, no soft-core, no vendor neural-net IP. A PC webcam streams 96×96 grayscale frames over UART; the FPGA classifies each frame as person or no person on-chip and reports the result back over UART. 83.3% accuracy on a 60-image hardware batch test, matching the 83.2% measured in quantized simulation before deployment — no meaningful hardware-specific degradation.
↗ github.com/Prateek2174/tang_primer_human_cnn- FPGA
- Sipeed Tang Primer 20K — Gowin GW2A-18C
- Input
- PC webcam → OpenCV → UART → 96×96 grayscale
- Network
- 4× Conv (3×3) + ReLU + MaxPool → GAP → FC (binary)
- Weights
- int8 QAT + BatchNorm fold + knowledge distillation → on-chip pROM
- Dataset
- INRIA Person (full scenes) · 614 images/class after balancing
- Accuracy
- 83.3% on real FPGA hardware · 83.2% quantized simulation
- BRAM
- ~32 of 46 BSRAM blocks used (binding resource)
- Toolchain
- Gowin EDA · PyTorch (training + distillation)
Context — sister project to the finger counter
The FPGA CNN accelerator proved out the core RTL architecture: a single reusable MAC datapath time-multiplexed across conv layers, int8 weights in pROM, and UART framing for deterministic input. This project applies that same architecture to a harder task — detecting human presence in uncropped real-world scenes — rather than building from scratch. The changes were in the depth of the network (four conv layers vs. three), the training pipeline (BatchNorm folding, knowledge distillation), and the dataset (INRIA Person, not self-collected).
What it does
A Python script on the PC captures a webcam frame, resizes it to 96×96
grayscale, and sends it over UART prefixed with the same
0xAA 0x55 sync marker used in the finger counter. The FPGA
assembles the frame, runs it through four convolutional layers, a global
average pooling stage, and a two-class fully connected classifier, then
sends the result — person (1) or no person (0) — back
as a single byte.
All weights are quantized to int8 and stored in on-chip pROM. Inference runs at 27 MHz (the board's raw oscillator, no PLL) with no off-chip memory.
Demo
Pipeline
Each stage feeds the next through on-chip BSRAM scratch buffers:
- UART RX + frame assembly — receives bytes at 115,200 baud, watches for the
0xAA 0x55sync marker, then writes the next 9,216 bytes into the input BSRAM, centering each pixel (pixel − 128) for signed CNN arithmetic. - CONV1 — 3×3 convolution over 96×96×1 → 96×96×12, then a 4×4 max-pool (stride 4) → 24×24×12. The 4×4 pool (instead of the usual 2×2) is a deliberate design choice: it halves the feature map again early, which shrinks BSRAM usage enough to fit the wider network on the 20K.
- CONV2 — 3×3 over 24×24×12 → 24×24×24, then 2×2 max-pool → 12×12×24.
- CONV3 — 3×3 over 12×12×24 → 12×12×48, then 2×2 max-pool → 6×6×48.
- CONV4 — 3×3 over 6×6×48 → 6×6×48, then 2×2 max-pool → 3×3×48.
- Global average pooling — averages each of the 48 channels in the 3×3×48 feature map to a single value, producing a 48-element vector.
- Classifier — a fully connected layer over the 48 GAP outputs; argmax of the two logits gives the binary prediction.
- Output — sends the class byte back to the PC over UART.
| Layer | Kernel | Stride | Padding | Channels | Pool | Output size |
|---|---|---|---|---|---|---|
conv1 | 3×3 | 1 | 1 | 1→12 | 4×4 | 24×24×12 |
conv2 | 3×3 | 1 | 1 | 12→24 | 2×2 | 12×12×24 |
conv3 | 3×3 | 1 | 1 | 24→48 | 2×2 | 6×6×48 |
conv4 | 3×3 | 1 | 1 | 48→48 | 2×2 | 3×3×48 |
| GAP | — | — | — | — | — | 48×1 |
| FC | — | — | — | 48→2 | — | 2×1 |
Training pipeline
Naively quantizing a float model to int8 after training lost too much accuracy on this task. The training pipeline went through three stages:
- Float training with real BatchNorm — train the network normally in PyTorch with BatchNorm layers after each conv. BatchNorm helps convergence but can't run in the RTL directly — it requires per-batch statistics that don't exist at inference time.
- BatchNorm folding — algebraically merge each conv layer's BatchNorm scale and shift into the conv weights and biases. After folding, the network is mathematically equivalent at float precision but has no BatchNorm nodes — it can be implemented as plain conv + bias + ReLU in hardware.
- QAT fine-tuning — fine-tune the folded model under real int8 quantization-aware arithmetic. This is what closes most of the gap between float and fixed-point accuracy.
On top of this, knowledge distillation was used: the small model being deployed was trained against both the ground-truth hard labels and the soft probability outputs of a larger teacher network. The teacher's soft outputs carry richer gradient signal than one-hot labels — the student learns not just which class is correct but how similar the model thinks classes are. This measurably improved quantized accuracy from 77.7% to 83.2% at zero extra hardware cost.
Dataset
The INRIA Person dataset — full uncropped scene photos, not tight crops around people. This is a harder problem than the finger counter: backgrounds are varied (streets, parking lots, buildings), lighting changes, and people appear at different scales and positions within the frame.
The raw dataset had a class imbalance. During training, tracking per-class accuracy revealed the model was collapsing — biasing toward the majority class to get easy loss reduction. The fix was to balance the dataset to 614 images per class before training, which resolved the collapse and let the model learn a real decision boundary.
Accuracy — simulation vs. hardware
The hardware accuracy (83.3%, measured over a 60-image batch test sent via UART to the real flashed board) is nearly identical to the quantized simulation accuracy (83.2%) run in PyTorch before deployment. This is the result that matters: it means the RTL faithfully implements the trained model — there is no hardware-specific degradation beyond what the int8 quantization itself already accounts for in simulation.
BRAM budget
BRAM is the binding resource on the GW2A-18C, not logic. The four feature map buffers (24×24×12, 12×12×24, 6×6×48, 3×3×48) use approximately 32 of the 46 available BSRAM blocks. The 4×4 pool on CONV1 was specifically chosen to control this: a standard 2×2 pool after CONV1 would produce a 48×48×12 feature map instead of 24×24×12, doubling the BSRAM requirement for that buffer and pushing the total past what the 20K can fit.
What was hard
The class-imbalance issue was the most time-consuming part of the training phase. A model that achieves 60%+ accuracy by always predicting the majority class is useless for detection — it took tracking per-class precision and recall separately (not just overall accuracy) to diagnose what was happening.
Knowledge distillation added training complexity: you need a pretrained teacher, a temperature-scaled softmax to spread the soft targets, and a loss that blends distillation loss with the hard-label cross-entropy. Getting the blend ratio right required experimentation. The jump from 77.7% to 83.2% justified the effort, but it was not a drop-in change.
On the RTL side, going from three to four conv layers with a wider channel count (up to 48 vs. 32 in the finger counter) required careful BSRAM partitioning. The 4×4 pool decision on CONV1 emerged from manually working through the BRAM budget before writing any RTL — verifying the fit in a spreadsheet before committing to the architecture.
What's next
- Live camera path — restore the OV5640 DVP input to eliminate the PC middleman; the UART path was chosen for debuggability, not as the final interface
- Multi-class extension — extend from binary (person / no person) to pedestrian counting or activity classification; will require a larger dataset and likely a channel-compression strategy to stay within the 20K's BSRAM budget
- On-chip display — drive an HDMI or SPI display from the FPGA showing the prediction result directly, without the UART return path to a PC