FPGA targets¶
Full path: Verilog/fpga/
Per-FPGA build and flow files. The Verilog/HDL source is shared and stays
under Verilog/ - the build scripts reference it by absolute path. Only the
board-specific build/flow files (scripts, constraints, tool projects) live here,
one folder per board.
Ready-built bitstreams are on the Releases page with step-by-step loading guides: QUICKSTART-nexys4ddr.md, QUICKSTART-tang-nano-20k.md, QUICKSTART-mister.md, QUICKSTART-mega65.md and QUICKSTART-qmtech-a35t.md. The MEGA65 cores and the QMTECH bitstream are built and timing-clean but have never run on their board - the first testers are the users, and both quickstarts say so and say what to report back. Release process: RELEASE-PLAN.md.
Targets¶
| Target | FPGA | Toolchain | Status | Details |
|---|---|---|---|---|
| tang-nano-20k/ | Gowin GW2AR-18 |
Gowin EDA (every variant incl. fast20; release builds) + OSS yosys/nextpnr (slow/crawl/full; built the full CPU 12-JUL-2026, but today maps it to 22,254-22,626 LUT4 of 20,736 and cannot place it - measured 28-SEP-2026, see ../TODO.md) |
Primary target - BOOTS SINTRAN III on silicon (24-AUG-2026), banner in 29.4 s from a Winchester image on the SD card. Full CPU bitstream with 4 MB SDRAM main memory (packed 16-bit storage + computed parity, ND_SDRAM_PACK16; other 4 MB reserved for the SD disk-image cache); SD/FAT stack proven on hardware (read+write, safety-gated) |
tang-nano-20k/README.md |
| basys3/ | Xilinx Artix-7 xc7a35tcpg236-1 |
Vivado (Windows host) | OPCOM boots on hardware (tag fpga-opcom-working-basys3); active debug line at 16.67 MHz; SD-card Pmod test build included (basys3/sd-fat-test/, Pmod JB). 24K-word memory ceiling (07-SEP-2026) - see cmod-a7-35t/README.md. |
basys3/README.md |
| cmod-a7-35t/ | Xilinx Artix-7 xc7a35t-1cpg236 (Digilent Cmod A7-35T DIP module, 512 KB external SRAM) |
Vivado (Windows host), same flow as Basys3 (same part) | First built 04-SEP-2026: fits easily (5,285 of 20,800 LUTs - a "11,493" figure here was WRONG, corrected 07-SEP-2026 from its own util.rpt) but MISSES TIMING, WNS -89.8 ns at 27 MHz - the CGA IDB ring cut at an unlucky point, not the board. 24K-word memory ceiling (07-SEP-2026) - see cmod-a7-35t/README.md. Cause of the 234 levels: the tool winds the ring FIVE times in one path (QMTECH winds it once); measured 07-SEP. (see ../../docs/HANDOFF-cga-idb-ring-cut.md); no bitstream written. Two build-script defects fixed on the way (missing include paths, runtime PROM-to-WCS load). 512 KB pack16 SRAM main-memory bridge planned (cmod-a7-35t/SRAM-BRIDGE-PLAN.md) |
cmod-a7-35t/README.md |
| nexys4ddr/ | Xilinx Artix-7 xc7a100tcsg324-1 (Digilent Nexys 4 DDR = Nexys A7-100T; 128 MiB DDR2, microSD, ~607 KB BRAM) |
Vivado (Windows host), Basys3 flow as template | SINTRAN III boots (25-AUG-2026), deployed at 33.333 MHz with the cache ON (31-AUG); 45.45 and 50 MHz also booted (cache off), 45.45 soaked 4 h; 27-AUG: SD-card deployment end to end (configure + boot from one card, no PC software) - DDR2-backed main RAM with BRAM cache; frequency search + bottlenecks in nexys4ddr/timing.md, boot record in nexys4ddr/SINTRAN-BOOT-25AUG.md, SD path in QUICKSTART-nexys4ddr.md |
nexys4ddr/README.md |
| qmtech-a35t/ | Xilinx Artix-7 xc7a35tcsg325-1 (QMTECH XC7A35T SDRAM core board, 32 MB SDRAM) |
Vivado (Windows host), same flow as Basys3 | BITSTREAM BUILT 04-SEP-2026: timing met, WNS +4.645 ns at 20 MHz, 12,619 of 20,800 LUTs, 22 of 50 BRAM tiles. Not yet run on the board. The 16 LUTLP-1 loops are downgraded to Warning to let write_bitstream run, as on the Nexys and MEGA65, so that slack is a floor not a guarantee. The whole machine: 4 MB SDRAM main memory via the sheet-49 bridge in 16-bit module mode (the MiSTer/MEGA65 configuration), SD-card storage and serial console on header JP3. The one Artix-7 target that can run SINTRAN - same die as the Basys3 without its 24 KB memory ceiling. Storage runs UNCACHED: the 16-bit bridge mode has no 32-bit access for the cache's region port. Wiring the console and SD Pmod to JP3, JTAG loading and the console test: QUICKSTART-qmtech-a35t.md |
qmtech-a35t/README.md |
| mega65/ | Xilinx Artix-7 xc7a200tfbg484-2 (MEGA65 retro computer; R3: 8 MiB HyperRAM, R4/R5/R6: + 64 MiB SDR SDRAM; 100 MHz osc) |
Vivado 2026.1 in-memory flow on the MiSTer2MEGA65 framework (git submodule) | Whole machine built for BOTH revisions (02-SEP-2026), timing-clean, not yet run on a MEGA65 - CPU + 4 MB main memory (R6: SDRAM via the MiSTer sheet-49 bridge, 20 MHz; R3: HyperRAM via the Nexys cache seam + an Avalon port, 13.33 MHz), TDV2200 terminal on the machine's own keyboard/screen (VGA + HDMI), floppy 0/1 + Winchester 0/1 + tape on the framework's virtual drives, one .cor per revision; every new block has a self-checking bench (mega65/docs/00-plan.md, QUICKSTART-mega65.md) |
mega65/README.md |
| mister/ | Intel Cyclone V SE 5CSEBA6U23I7 (DE10-Nano / "MiSTer PI", ~110K LE + ARM HPS running Linux) |
Quartus Lite 17.0.2 (free, Docker raetro/quartus:17.0) |
BOOTS SINTRAN III on silicon (02-SEP-2026) - the whole ND-120 machine: 4 MB main memory in the DE10-Nano SDRAM module, TDV2200 console on the MiSTer's own screen + keyboard, floppy 0/1 + Winchester 0/1 + tape as Linux-side image files from the OSD, CPU at 20 MHz. Hardware-verified .rbf in Release 2; load guide QUICKSTART-mister.md |
mister/README.md |
| azure-x930613/ | Intel/Altera Stratix V GS (Microsoft Azure FPGA 40GbE QSFP+ PCIe, P/N X930613-001, PCB DAT6MTHUEB0; 4 GB DDR3 ECC, 2x QSFP+/40GbE, PCIe Gen3 x16) | Quartus Prime Standard (per third-party report, unverified against a datasheet) | Paper plan (31-AUG-2026) - a datacenter PCIe accelerator card, not a devboard: no confirmed console/GPIO, custom non-catalog FPGA part per a third-party blog, real specs still unverified against a datasheet or the physical card | azure-x930613/README.md |
How fast can each device run the CPU - and what stops it¶
The honest yardstick is the machine's one critical-path family, identical on
every target: the full microcycle - WCS microcode BRAM ->
CSIDBS/INTR/ALU/FIDBO/TRAP/PALs/MIC/ACAL -> the WRF register-file
clock-enables, ~30 logic levels of combinational PAL/TTL transcription with
no register in between. Whoever executes those ~30 levels fastest wins.
Full analysis: nexys4ddr/timing.md and
nexys4ddr/timing-analysis/TIMING_CLOSURE_REPORT.md.
Legend: measured = read from that board's own post-route timing report or proven on its silicon; estimate = transferred from a measured board with the same die/fabric; unknown = never built or never measured.
| Device | Fabric | Microcycle on this fabric | Honest CPU ceiling (STA) | Proven on silicon | What actually limits it |
|---|---|---|---|---|---|
Nexys 4 DDR xc7a100t-1 |
28 nm Artix-7, LUT6 | measured: 29-30 levels, ~0.73 ns/level, 21.7 ns total at the wall | measured: 45.45 MHz default flow, 50 MHz with phys_opt (both single-seed; 50 is fragile) |
deployed at 33.333 MHz cache ON; SINTRAN also at 45.45 MHz (cache off, 4 h soak 8/8) and 50 MHz | Nothing structural left below ~45 MHz. Beyond: the microcycle itself (only a pipeline breaks it, which kills cycle-faithfulness). DDR2 is decoupled in its own 75 MHz domain, so memory never gates the CPU clock |
Tang Nano 20K GW2AR-18 |
~55 nm Gowin Arora, LUT4 (vendor data) | measured: 32 levels, ~1.5 ns/level, ~49 ns total (tang-nano-20k/build/.../nd120_tang20k_build.tr) |
measured: Actual Fmax 22.932 MHz (fast20, after the 31-AUG .sdc CDC fix; was 20.6) |
SINTRAN at 20.25 MHz + 115200 console, TIMING-CLEAN (TNS 0), booted 26-AUG-2026, 4 h soak 8/8 probes - the fast20 variant, 3x the long-validated 6.75 MHz. 27 MHz also boots but runs 32% past its own Fmax (1667 violations, margin unquantified) |
1) fabric ~2x slower per level than Artix-7 (physics), 2) Gowin flow has no phys_opt and no WNS gate, 3) SDRAM clocks share the rPLL VCO with the CPU clock (cap ~33 MHz), 4) the .sdc was one line - FIXED 31-AUG-2026, see tang-nano-20k/README.md. Bottlenecks 1-3 stand |
Basys3 xc7a35t-1 |
same 28 nm Artix-7 fabric as the Nexys | estimate: identical per-level speed (same die family, same speed grade) | last measured 21-AUG-2026: WNS -29.8 at 16.667 MHz - PRE-ring-cut and stale; never re-measured after commit b3ee391 cut the FIDBO ring |
OPCOM boots; SINTRAN impossible regardless of clock | Capacity, not speed: 100 RAMB18 -> 24 KB main RAM config. The fabric could do Nexys-class clocks; there is no memory to run an OS in |
Cmod A7-35T xc7a35t-1 |
same 28 nm Artix-7 fabric | estimate: identical per-level speed | first build 04-SEP-2026 misses timing (WNS -89.8 ns at 27 MHz), no bitstream written | not run - no bitstream | Capacity until the 512 KB SRAM bridge lands (cmod-a7-35t/SRAM-BRIDGE-PLAN.md); then the SRAM protocol timing becomes the question, not the fabric |
QMTECH XC7A35T xc7a35t-1 |
same 28 nm Artix-7 fabric | measured 04-SEP-2026, first build: CPU domain closes at 20 MHz with +5.255 ns slack over 27,698 endpoints, none failing | headroom exists but is unmeasured - 20 MHz was the target, not the ceiling | not yet run on the board | Bitstream built, timing met at +4.645 ns. 12,619 of 20,800 LUTs, 22 of 50 BRAM tiles. The CGA IDB ring is present (16 LUTLP-1, 10 auto-cuts) and lands harmlessly on this netlist - unlike the Cmod, same die family, where it costs 89.8 ns. The SDRAM bridge caps its own 2x clock at 66.7 MHz, so 20-33 MHz is the band |
MiSTer / DE10-Nano 5CSEBA6U23I7 |
28 nm Cyclone V SE, ALM (vendor data) | synthesized; per-level figure not separately reported | not separately measured (runs at 20 MHz) | SINTRAN at 20 MHz + TDV console, deployed 02-SEP-2026 | Runs at 20 MHz; the fabric ceiling is not yet measured. 28 nm ALM should land between the Tang and the Artix-7 per level - an inference until a Quartus timing report says otherwise |
MEGA65 xc7a200t-2 |
same 28 nm Artix-7 fabric as the Nexys, one speed grade faster | measured 02-SEP-2026 on the post-route checkpoints: the WCS -> MAC microcycle path is 34 ns in the R6 netlist (58 levels) and 57 ns in the R3 netlist (93 levels) - the R3 netlist times the CGA IDB ring through a longer loop-break point, the same "impossible but unprovable" ring the Nexys documents | measured: R6 closes at 20 MHz with 15.9 ns of CPU-domain slack; R3 closes at 13.33 MHz with 29.8 ns | nothing yet - no MEGA65 here; the first testers are the users of the release cores | The IDB ring, not the fabric: the period was set to fit the ring path rather than untime the internal data bus. Cut the ring in RTL and the R3 goes to 20 MHz+ like the R6; the R6 has room for ~40 MHz on paper (unmeasured) |
Three portable lessons from the Nexys campaign that apply to every row:
- A loose constraint hides the real ceiling. At 60 ns the Nexys microcycle "took" 33 ns; under pressure the router compressed the same logic to 21.7 ns. Estimating Fmax from a relaxed run's WNS undershot the demonstrated ceiling by 50%.
- A WNS gate is what makes numbers mean anything. The Vivado flows refuse to write a bitstream over negative slack; the Gowin flow does not, which is how the Tang full (27 MHz) variant boots with 1667 violations; the released Tang file is fast20, TNS 0 - on margin nobody has quantified.
- Closures at the wall are single-seed lottery tickets. Changing one UART divider constant re-rolled the Nexys 50 MHz closure from +0.007 ns to -0.210 ns FAIL. Near the wall, every edit needs its own clean report.
Board status and priority (28-SEP-2026)¶
- Tang Nano 20K - primary target: boots SINTRAN (
fast20, 20.25 MHz, timing-clean). It is also the project's value-for-money benchmark: under 300 NOK for SDRAM + microSD + USB-JTAG/UART + HDMI - judge any new board suggestion against it. - Nexys 4 DDR: boots SINTRAN, deployed at 33.333 MHz with the cache ON.
- MiSTer (DE10-Nano): boots SINTRAN (02-SEP-2026).
- QMTECH XC7A35T - the SINTRAN-capable Artix-7 target: built
04-SEP-2026, timing met, not yet run on the board. 4 MB of main memory,
which is the whole of the ND-120's onboard memory space - the CPU board
decodes onboard memory in the bottom 2M words (
PAL_44445B.v:85), so the other 28 MB of the chip could not be addressed as main store however it were wired. - MEGA65: built for both revisions (02-SEP-2026), timing-clean, not yet run on a MEGA65.
- Cmod A7-35T: first build misses timing (WNS -89.8 ns at 27 MHz), no
bitstream; 24K-word memory ceiling until the 512 KB SRAM bridge lands
(
cmod-a7-35t/SRAM-BRIDGE-PLAN.md). - Basys3: OPCOM booted on hardware (07-JUL-2026); the 24K-word memory ceiling rules out SINTRAN, so it stays the ILA/debug board.
Per-board detail lives with the board (README, vendor docs, plans and
handoffs in each <board>/ folder) - this file is only the directory. For
the QMTECH, see qmtech-a35t/README.md.
Prerequisites¶
Every tool needed - Vivado, Gowin EDA, Verilator, iverilog, yosys, the
serial/JTAG access route, and the documentation generators - with exact
install and validation commands and the versions in use, is documented in
../docs/PREREQUISITES.md.
Building - one API for every board¶
Every board folder has a Makefile with the same targets, whatever the
toolchain underneath. From WSL (the Windows-hosted tools are reached via
powershell.exe - works because the repo lives on a Windows drive):
| Target | Meaning |
|---|---|
make |
Build the bitstream (no board needed) |
make load |
Program the FPGA - volatile (JTAG/SRAM; gone at power-cycle) |
make flash |
Program persistent config flash (survives power-cycle) |
make sim |
Run the board folder's iverilog testbenches (where present) |
make clean |
Remove build outputs (where present) |
Board-specific extras: basys3 adds make reuse (skip the ~1h resynth,
reuse the synth_1 checkpoint) and make lint; qmtech-a35t takes
TEST=led-test|mem-test (default mem-test) and has no flash flow yet;
nexys4ddr takes CLK=<MHz> (default 16) plus CACHE=0 / VGACONSOLE=0 /
PANELCLOCK=0, and has the full load (JTAG, volatile) / flash (QSPI,
permanent) pair;
mister has a Quartus project (DE10-Nano / Cyclone V) and builds via the
Quartus-in-Docker flow - make build / make load / make flash / make sim;
its console is the TDV2200 terminal, same as the Nexys; cmod-a7-35t
uses the same Vivado flow as Basys3 (make / make build / make clean,
delegating to build.tcl). The standalone
tang-nano-20k/sdram-test/ keeps its own Makefile with the same
all/load/flash/sim/clean targets (Linux-native OSS flow).
On the Windows host the underlying scripts are the direct entry points (the Makefiles just delegate to them):
cd Verilog/fpga/basys3
.\vivado_build.ps1 # bitstream ('make'); -ReuseSynth / -LintOnly / -Program
.\flash.ps1 -Quick # 'make load' (JTAG only); omit -Quick for 'make flash'
cd ..\tang-nano-20k
.\gowin_build.ps1 # bitstream ('make'); copies WCS preload + checks EX3988
cd ..\qmtech-a35t\mem-test # or led-test
vivado -mode batch -source build.tcl -tclargs skip_program # 'make'
vivado -mode batch -source build.tcl # 'make load'
Programming transport per board: Basys3 and Cmod A7 = onboard USB-JTAG;
Tang Nano 20K = openFPGALoader from WSL (usbipd-attached) or the Gowin
programmer GUI; QMTECH = Xilinx Platform Cable USB II on the JTAG header.
Local settings: run configure.py once¶
No build script names a folder on anybody's machine. Paths inside the repo
are worked out from each script's own location; paths outside it are
settings in local.mk at the repository root, which
python3 configure.py (py configure.py from a Windows shell) writes: it
finds the tools, asks for what it cannot find, and remembers the answers. The
table of every setting - what it is, what reads it, whether it is required -
is in CONTRIBUTING.md - Local settings;
local.mk.example explains each one too. The board
builds need ND120_BUILD_DIR (where builds go) and the vendor tool
(ND120_VIVADO, or ND120_GOWIN for the Tang's Gowin flow).
Every board Makefile includes paths.mk at the repository
root, which loads local.mk and exports the values (also through WSLENV,
so powershell.exe / cmd.exe started from WSL see them). The PowerShell
scripts read local.mk through paths.ps1 and the Tcl scripts
through paths.tcl, so running a script by hand needs nothing
more. A value in the environment wins over local.mk; make VIVADO=...
still overrides the Vivado path for one run. Each target checks the settings
it needs before doing any work and names what is missing; make
check-config in any board folder shows them all.
Where builds go: every board writes everything - bitstream, reports,
timing-analysis runs, checkpoints, copied microcode, Vivado's log, journal and
.Xil - to $ND120_BUILD_DIR/<board>/, never into this tree. make
fresh-build ND120_FRESH_DIR=<empty folder> in a board folder clones the
current commit there, configures it with a build folder inside the clone and
builds it - the proof that a build needs nothing outside the repository.
Shared context (applies to all boards)¶
- FF mode (single
sysclk+ clock-enables) is what lets the boards boot: Tang, Nexys and MiSTer run SINTRAN. See../docs/fpga-debug-methodology.md3.2. - Microcode preload:
SKIP_WCS_LOADbitstream-preloads the WCS and skips the runtime load phase - verified in Verilator, and required to fit the Tang's BSRAM. See../docs/skip-wcs-load.md. - Compile-time defines (per-target behavior): see
../docs/build-defines.md. - Golden boot reference for validation: see
../docs/boot-golden-spec.md.
Reference docs¶
../docs/fpga-debug-methodology.md- Verilator-vs-FPGA debug../docs/build-defines.md- compile-time defines../docs/skip-wcs-load.md- preloaded-WCS microcode../docs/boot-golden-spec.md- expected boot sequence../docs/basys3-memory-speed-validation.md- which memory backends meet the no-wait-state protocol (per board)../sim/FPGA_DEBUG_RUNBOOK.md- Verilator-vs-board comparison method