case 12 · Personal · 2026

FPGA CNN Accelerator — Verilog hand gesture classifier

A three-layer convolutional neural network built entirely in Verilog on a Tang Primer 20K — no CPU, no soft-core, no vendor neural-net IP. A PC webcam streams 96×96 grayscale frames over UART; the FPGA classifies the hand gesture (0–5 fingers) on-chip and reports the result back over UART and on the board LEDs. 84.3% accuracy on 300 real hardware test images, against an 88.1% validation accuracy of the trained model.

Verilog FPGA Tang Primer 20K CNN int8 UART Gowin EDA
↗ github.com/Prateek2174/tang-primer-cnn
FPGA
Sipeed Tang Primer 20K — Gowin GW2A-18C
Input
PC webcam → OpenCV → UART → 96×96 grayscale
Network
3× Conv (3×3) + ReLU + 2×2 MaxPool → GAP → FC
Weights
int8 quantization-aware training → on-chip pROM
Dataset
3,547 self-collected images · 6 classes (0–5 fingers)
Accuracy
84.3% measured on FPGA · 88.1% validation (PyTorch)
Logic
9% LUT/ALU/ROM16 · 6% registers · 59% BRAM (binding)
Toolchain
Gowin EDA V1.9.11 · PyTorch (training)

What it does

The FPGA runs a complete CNN inference pipeline — from raw pixel input to classified output — with no processor of any kind involved. A Python script on the PC captures a webcam frame, resizes it to 96×96 grayscale, and sends it over UART prefixed with a 0xAA 0x55 sync marker. The FPGA assembles the frame, runs it through three convolutional layers, a global average pooling stage, and a fully connected classifier, then sends the predicted finger count (0–5) back as a single byte and lights the corresponding LED.

All weights — the convolutional filters, per-filter biases, and FC layer weights — are quantized to int8 and stored in on-chip pROM. The entire inference runs in a single clock domain at 27 MHz (the board's raw oscillator, no PLL) with no off-chip memory.

Pipeline

Each stage feeds the next through on-chip BSRAM scratch buffers:

  • UART RX + frame assembly — uart.v receives bytes at 115200 baud. uart_frame.v watches for the 0xAA 0x55 sync marker, then writes the next 9,216 bytes into the resize BSRAM, centering each pixel (pixel − 128) for signed CNN arithmetic.
  • CONV1 — 3×3 convolution over 96×96×1 → 48×48×8, folding bias-add, ReLU, and 2×2 max-pool into the same MAC FSM. Output stored in Feature Map A (BSRAM).
  • CONV2 — 3×3 over 48×48×8 → 24×24×16, same structure. Output in Feature Map B.
  • CONV3 — 3×3 over 24×24×16 → 12×12×32. Output in Feature Map C.
  • Global average pooling — global_avg_pool.v averages each of the 32 channels in Feature Map C to a single value, producing a 32-element vector.
  • Classifier — classifier.v runs a fully connected layer over the 32 GAP outputs and takes the argmax to produce the finger-count class.
  • Output — uart_tx.v sends the class byte back to the PC; the top-level also drives one LED per class in one-hot encoding.
Architecture / conv hyperparameters
LayerKernelStridePaddingChannelsBias
conv13×3111→8yes
conv23×3118→16yes
conv33×31116→32yes
max pool (after each conv+ReLU)2×220——
FC———32→6no

MAC array design

Rather than instantiating three separate conv engines, I wrote a single reusable MAC datapath in mac_array.v and time-multiplexed it across all three conv layers via a conv_layer_sel signal. The MAC FSM folds bias-add, ReLU, and 2×2 max-pool into its own state machine — there are no separate pool or ReLU modules in the live pipeline. conv_acc.v implements the inner 9-tap signed multiply-accumulate for one conv window; mac_array.v loops over channels and spatial positions and coordinates writes back to the feature map BSRAMs.

BRAM (59% utilised) is the binding resource, not logic — the three feature maps (48×48×8, 24×24×16, 12×12×32) together require 27 of the 46 available BSRAM blocks.

Weights and quantization

The model was trained in PyTorch with quantization-aware training targeting int8 weights. After training, weights are extracted, quantized, and written into Gowin pROM initialization files. weight_rom.v exposes them as an addressed read interface to the MAC array. Biases are also int8, stored per filter. The ~4 percentage-point drop from validation accuracy (88.1%) to real hardware accuracy (84.3%) is accounted for by the fixed-point rounding in the RTL, confirmed by running the same int8 inference in PyTorch and getting matching numbers — no hardware bug.

Dataset

3,547 images collected by hand — six classes (0 to 5 fingers), shot against varied backgrounds and lighting, multiple subjects. 85/15 train/val split (3,017 / 530 images). The dataset collection was the longest phase of the project; getting enough variation in background and hand orientation to make the model generalize was what made the final 84.3% on real hardware hold up.

Demo

Demo 1 — live classification running on the FPGA.
Demo 2.
Demo 3.
Demo 4.
Demo 5.

Design notes

Working notes from the design phase — the CNN datapath structure, BSRAM addressing, and MAC FSM carried forward into the final build.

Handwritten notes showing ReLU, MaxPool logic and CONV1→CONV2→CONV3 channel and dimension flow.
ReLU/MaxPool logic and the dimension flow: 96×96×1 → 48×48×8 → 24×24×16 → 12×12×32 — unchanged from the sketch to the final RTL.
Handwritten notes showing feature map BSRAM layout for the three conv layers.
Feature map BSRAM layout — 48×48×8, 24×24×16, 12×12×32 and the conv core/accumulator split.
Refined MAC array FSM sketch with per-channel accumulate loop and pool-index increment logic.
Refined mac_array.v FSM — adds the per-channel accumulate loop and pool-index increment that made it into the final implementation.
Handwritten notes showing weight ROM byte counts and base addresses per conv layer.
Weight ROM sizing — byte counts and base addresses per conv layer (CONV1/2/3), used to generate the pROM initialization files.

What was hard

The biggest challenge was fitting the feature maps in BRAM. At 59% utilisation BRAM was the binding resource — logic sat at 9%. The three feature map buffers (48×48×8, 24×24×16, 12×12×32) together consume 27 of the 46 available BSRAM blocks. Getting the addressing right — the conv window slides across two spatial dimensions and a channel dimension simultaneously — required careful indexing math to avoid BSRAM port conflicts and to keep the MAC loop on schedule.

The earlier design fed a live OV5640 camera directly into the FPGA over a DVP parallel interface, but that path had frame-sync and decimation complexity that was adding noise to the debugging. Switching to UART from a PC webcam gave full control over the input — exact pixel values, known frame boundaries, reproducible test images — which is what made the cycle-accurate simulation vs. hardware accuracy comparison possible.

Quantization-aware training was necessary. Naively quantizing a float32-trained model to int8 post-hoc dropped accuracy below 70%. Retraining with QAT and re-quantizing recovered it to 88.1% val / 84.3% hardware.

What's next

  • Live camera input — restore the OV5640 DVP path now that the CNN datapath is stable; eliminates the PC middleman and makes the system fully self-contained
  • Pipeline parallelism — the current design processes one conv layer at a time sequentially; overlapping CONV2 with CONV1's output write-back would cut latency
  • More classes — extend beyond 0–5 fingers to ASL letters or other gestures; requires a larger dataset and more filters (wider network), which will push BRAM usage past 100% and force a CONV3 feature map compression strategy
  • On-chip display output — drive an HDMI or SPI display directly from the FPGA showing the classified gesture alongside the confidence scores, without needing the UART return path to a PC