Tang Nano 20K (Gowin GW2AR-18) FPGA target¶
Full path: Verilog/fpga/tang-nano-20k/
Status (02-SEP-2026): SINTRAN III BOOTS ON REAL SILICON. The Tang Nano 20K was the FIRST board to boot SINTRAN (24-AUG-2026) and is a working machine you log into.
fast20runs it at 20.25 MHz, 115200 7E1 console, timing-clean (TNS 0); 6.75 MHz is the long-validated safe speed andmid(13.5 MHz) also closes. Main memory is 4 MB of the embedded SDRAM (packed 16-bit,ND_SDRAM_PACK16); the SD/FAT storage stack is proven on hardware. From the Winchester image on the SD card it reaches the SINTRAN banner in 29.4 s from cold; login,LIST-FILESand the S3 program work (S3 cold start 13.2 s).fast20is a Gowin EDA build only - the OSSMakefileoffersslow,crawlandfull. The old page-fault / silicon-hang / level-14 livelock / bank-decode campaigns are all RESOLVED - that is why it boots. Details below.
Gowin build/flow for the Sipeed Tang Nano 20K. This is the primary FPGA
target going forward - chosen for faster synthesis than Vivado, a Linux-native
open-source toolchain option, and 8 MB of SDRAM that lets the FPGA run the full
memory config like the simulator. (The pre-port analysis and staged plan,
Verilog/docs/tang-nano-20k-port.md, is finished and was deleted 28-SEP-2026;
git history keeps it.)
Bring-up: getting the board talking (do this first)¶
Every session starts here. The whole USB dance is already scripted - it does not need to be done by hand, and doing it by hand is how most of the wasted time on this board has been spent.
cd Verilog/fpga/tang-nano-20k
make usb # find the Tang on the Windows host, attach it to WSL,
# open the raw USB node AND /dev/ttyUSB* permissions
make flash-gowin # program the current bitstream
After make usb you have:
| Device | Use |
|---|---|
/dev/ttyUSB0 |
JTAG side (openFPGALoader) |
/dev/ttyUSB1 |
OPCOM console, 115200 7E1 |
Reflashing reboots the design - no power cycle needed. Measured
09-AUG-2026: make flash-gowin restarts the ND-120 and the console comes back
to the # prompt on its own. So a build/flash/test cycle can run unattended.
A real power cycle is different. Pulling power detaches the device from
WSL and the host busid CHANGES (seen 3-3, then 2-4, then 5-2 on different
days). After a power cycle, run make usb again - never assume the old busid.
When it will not talk¶
| Symptom | Cause | Fix |
|---|---|---|
unable to open ftdi device: -4 (usb_open() failed) |
the raw USB node's permissions reset when the device re-enumerated | make usb, or sudo chmod 666 $(lsusb -d 0403:6010 \| sed -E 's#Bus ([0-9]+) Device ([0-9]+).*#/dev/bus/usb/\1/\2#') |
unable to open ftdi device: -3 (device not found) |
not attached to WSL at all | make usb |
cannot open /dev/ttyUSB1: No such file or directory |
attached, but ftdi_sio has not created the serial nodes |
make usb |
openFPGALoader: command not found |
running the command on the Windows side instead of inside WSL | run it from WSL |
Driving the console¶
OPCOM is picky and the rules are not guessable:
- UPPERCASE only.
- ~0.30 s between characters. At 0.12 s characters are dropped silently in
the middle of a number and OPCOM answers
?- which looks like a machine fault and is not one. - 115200 7E1 on
/dev/ttyUSB1. Since 26-AUG-2026UART_BAUD_RATEis 115200 for EVERY variant (src/tang20k_defines.v:554, unconditional) - scripts written before that opened 9600 and must be updated.
Use the committed driver rather than writing another one:
Command reference: https://nd110.hackercorp.no/Terminal; boot commands in
NorskData-Doc/OPCOM-Boot-Reference.md.
Board / device¶
Hardware reference: Sipeed wiki - Tang Nano 20K (datasheets/schematics: dl.sipeed.com, official examples: sipeed/TangNano-20K-example).
| Item | Value |
|---|---|
| Board | Sipeed Tang Nano 20K (22.55 mm x 54.04 mm) |
| FPGA | Gowin GW2AR-LV18QN88C8/I7 (GW2AR-18, QN88) |
| Logic | 20,736 LUT4, 15,552 FF, 48 18x18 multipliers, 8 I/O banks |
| Block RAM (BSRAM) | 828 Kbit (46 blocks) + 41,472 bit shadow SRAM |
| Big RAM | 8 MB SDRAM (64 Mbit, 32-bit SDR, embedded in the GW2AR package; no package pins - see sdram-test/) |
| Config flash | 64 Mbit |
| Clock | 27 MHz crystal + MS5351 clock chip (3 extra clocks, BL616-controlled), 2 PLLs (use a Gowin rPLL for the CPU clock) |
| Programmer | BL616 onboard USB (JTAG + USB-UART + USB-SPI + MS5351 control); openFPGALoader-compatible; enumerates as one USB serial port (COM5 on Ronny's machine) |
| Peripherals | HDMI, 40-pin RGB LCD connector, TF-card slot, MAX98357A audio amp |
| UI | 6 LEDs (active low, pins 15-20), 1 WS2812 RGB LED, 2 buttons (S1 = pin 88, S2 = pin 87) |
Factory firmware: the flash ships with a LiteX SoC (VexRiscv_Min @ 48 MHz,
from the litex/ folder of the TangNano-20K-example repo; prebuilt
litex/tang_nano_20k_litex.fs). Its BIOS talks at 115200 baud and has
mem_test / mem_speed / sdram_test commands - an independent check that the
SDRAM hardware works. Its boot report confirms the SDRAM at 48 MHz CL-2 with a
clean 2 MB memtest, mapped at 0x40000000 (8 MB). Loading a bitstream into SRAM
(openFPGALoader without -f) leaves it intact; flashing replaces it (restore with
openFPGALoader -b tangnano20k -f .../litex/tang_nano_20k_litex.fs).
vs Basys3¶
Fewer LUTs (20,736 LUT4 vs 33,280 LUT6) and less BSRAM (828 vs ~1,800 Kbit), but adds 8 MB SDRAM and a much faster, Linux-friendly toolchain. Fit is the top risk (see Memory architecture).
Toolchain (two options)¶
Option 1 - Gowin EDA (authoritative)¶
gw_sh (Tcl, scriptable, driven by gowin_build.ps1) or the GUI, using
nd120_tang20k.gprj. Uses a Synplify-based synth that tolerates the design's TTL-style
flip-flops (clock + async preset + async clear). Has GAO (Gowin Analyzer
Oscilloscope), the on-chip logic analyzer = Vivado ILA equivalent. This is the
reliable path to a real bitstream today.
Local settings¶
No install path is written in the scripts. Run python3 configure.py at the
repository root once: it finds gw_sh.exe and writes ND120_GOWIN and the
build folder ND120_BUILD_DIR to local.mk (see
CONTRIBUTING.md - Local settings).
make gowin passes the settings on (through WSLENV from WSL), and
gowin_build.ps1 also reads local.mk itself when run by hand. The OSS flow
takes its tools from ND120_OSS_CAD_SUITE, else ~/oss-cad-suite, else PATH
(make OSS_CAD=... overrides for one run).
Output: both flows write everything to $ND120_BUILD_DIR/tang-nano-20k/
(the Gowin project and its impl/ reports, the OSS netlists and logs, the
copied WCS images) - nothing lands in this folder, and every build target,
make check included, stops with a message when ND120_BUILD_DIR is not set.
make fresh-build ND120_FRESH_DIR=<empty folder> runs make gowin on the
current commit in a fresh clone.
Option 2 - OSS flow (Linux-native, WSL)¶
yosys synth_gowin -> nextpnr-himbaechel --device GW2AR-LV18QN88C8/I7 ->
gowin_pack -> openFPGALoader. Runs entirely on WSL/Linux (no Windows context
switch).
Install (oss-cad-suite - the whole chain in one prebuilt bundle, no sudo):
cd ~
# resolve the latest linux-x64 tarball via the GitHub API, then extract
URL=$(curl -s https://api.github.com/repos/YosysHQ/oss-cad-suite-build/releases/latest \
| grep -oE '"browser_download_url":[^,]*linux-x64[^"]*\.tgz' | grep -oE 'https://[^"]*')
curl -L -o oss-cad-suite.tgz "$URL"
tar xzf oss-cad-suite.tgz # -> ~/oss-cad-suite/
# activate in each shell that uses the tools:
source ~/oss-cad-suite/environment
yosys --version # expect 0.4x+ (not the distro's 0.9)
which nextpnr-himbaechel gowin_pack openFPGALoader
pip install yowasp-yosys.)
The full CPU built with this flow on 12-JUL-2026, but it no longer fits: see Two build flows below. Use Gowin EDA for a bitstream.
OSS flow + embedded SDRAM: nextpnr does not auto-connect the on-package
SDRAM the way Gowin EDA does with the magic O_sdram_*/IO_sdram_dq port names -
they must be pinned explicitly to internal pseudo-pins (IOL13A etc.). The pin
list (from Seyviour/sdram-tang-nano-20k-os-example)
is vendored at sdram-test/src/sdram_pins_oss.cst and must be kept out of
the Gowin EDA constraints.
Two build flows¶
Everything runs from Verilog/fpga/tang-nano-20k/. Both flows compile the
same ordered file list parsed out of nd120_tang20k.gprj, with
src/tang20k_defines.v first, so every `define comes from one place (yosys
keeps defines across the files of one read_verilog, like Gowin's ordered
compilation unit). The clock variant is chosen without editing a file:
the OSS Makefile passes -DTANG_VARIANT_*; gowin_build.ps1 -Variant ...
writes build/tang20k_variant.v as the first project file on every build.
No variant define = slow.
| Gowin EDA | OSS CAD Suite | |
|---|---|---|
| Runs on | Windows host (gw_sh) |
WSL / Linux (ND120_OSS_CAD_SUITE, else ~/oss-cad-suite; override OSS_CAD=) |
| Build | .\gowin_build.ps1 [-Variant slow\|crawl\|mid\|full\|fast20] or make gowin |
make [VARIANT=slow\|crawl\|full] (no mid/fast20) |
| Bitstream | <build>/nd120_tang20k_build/impl/pnr/nd120_tang20k_build.fs |
<build>/nd120_tang20k_oss-<variant>.fs |
| Load / flash | make load-gowin / make flash-gowin |
make load / make flash |
| Netlist gates | EX3988 empty-WCS check in the ps1 | make check (IO_sdram_dq tristate + latch census; also in CI) |
What the OSS flow needed (all under `ifdef YOSYS, other flows untouched):
CPU_PROC_32.vregisterBlockselects distributed LUT RAM (yosys treatsram_style="block"as a hard requirement, and the read port is async).Shared/ndlib/SCAN_WITH_SET_N_EN.vpowers up to 1: yosysdfflegalizerejects a power-up value that differs from the async-set value.- nextpnr does not auto-connect the SDRAM ports: the Makefile appends
sdram-test/src/sdram_pins_oss.cstto the board .cst intobuild/oss.cst. - The WCS preload hex files resolve against the yosys working directory;
make wcs-hexrefreshes them. A missing file is a hard error. - nextpnr takes no SDC: judge timing by the Fmax in
...-pnr.log. First OSS builds (12-JUL-2026): clk_cpu Fmax 47.98 (slow), 48.99 (crawl), 57.51 MHz (full, 27 MHz needed). - On the Windows drive a WSL compile now and then misses a just-written file ("No such file or directory"); re-run.
The full CPU no longer fits with the OSS flow (measured 28-SEP-2026 with
the oss-cad-suite CI downloads): yosys maps it to 22254-22626 LUT4 against
20736 on the chip (107-109%), so nextpnr cannot place it - older suites sit in
the placer until they are stopped (CI run 33664876050: 2 h), newer ones stop
at once with no BELs remaining. No timeout fixes that. Gowin EDA fits the
same design (it maps it differently), so bitstreams come from Gowin.
Synthesis-setting work that brings it to 19790 LUT4 but still does not place
is recorded in Verilog/TODO.md (the OSS-fit item). No bitstream is built in
CI.
Full ND-120 build¶
The complete ND-120 CPU has a Tang top-level and Gowin project here:
| File | Purpose |
|---|---|
src/ND120_TANG20K_TOP.v |
Board top: instantiates ND3202D, ties off the external bus, S1 = Master Clear, OPCOM UART 115200 on the BL616 (pins 69/70), 6 LEDs: block-read/write activity, tape byte served, SD status pair, heartbeat (see the LED table below) |
src/tang20k_defines.v |
Must stay FIRST in the project - defines GOWIN, TARGET_TANG20K, FPGA_FF_MODE, SKIP_WCS_LOAD, MAIN_RAM_SDRAM, BOARD_CLK_FREQ (per clock variant), UART_BAUD_RATE=115200 |
src/gowin_rpll_27_54.v |
One rPLL: 54 MHz (SDRAM ctrl) + 54 MHz shifted (SDRAM chip) + 27 MHz (CPU/bus/OSC) |
src/nd120_tang20k.cst / .sdc |
Pins (verified 20K pinout); the 27 MHz input clock plus three reasoned set_false_path exceptions (see Clock variants, 31-AUG and 01-SEP notes) |
nd120_tang20k.gprj |
Gowin GUI project - 247 files, generated from the Verilator dependency list (single source of truth for the tcl too) |
gowin_build.tcl / gowin_build.ps1 |
Scripted build on the Windows host: .\gowin_build.ps1 copies the 32 WCS preload hex files (Code/Microcode/wcs/), runs gw_sh from the Gowin EDA install (ND120_GOWIN in local.mk at the repository root - see Local settings) -> <build>\nd120_tang20k_build\impl\pnr\nd120_tang20k_build.fs, <build> = $ND120_BUILD_DIR\tang-nano-20k |
lint/rpll_stub.v |
Lint-only rPLL stub (Verilator elaboration check; not in the Gowin build) |
Build config: microcode is bitstream-preloaded (SKIP_WCS_LOAD, PROM never
read), main memory is the 8 MB embedded SDRAM through
sdram-bridge/ (2 banks = 4 MB), CPU clock chosen by
the variant (below; slow = 6.75 MHz is the default). The Verilator sim build is
unaffected by the Tang defines.
Which toolchain builds it: Gowin EDA. The OSS flow (yosys/nextpnr) built the full CPU locally on 12-JUL-2026 (all three variants; Fmax in Two build flows) but no longer fits it (see there). Release bitstreams are Gowin EDA builds, made locally and checked on the board. The Fmax and violation figures in this README are Gowin EDA timing reports.
Clock variants and measured boot timings (24-AUG-2026)¶
Clock variants, selected with gowin_build.ps1 -Variant
<slow|mid|fast20|full>. clk_cpu is always exactly half of clk2x;
slow/mid/full share one 864 MHz VCO, fast20 uses 648 MHz.
| Variant | CPU | SDRAM | Setup violations | CPU-domain Fmax (Gowin STA) |
|---|---|---|---|---|
crawl |
3.375 MHz | 6.75 MHz | - | - |
slow (default) |
6.75 MHz | 13.5 MHz | 0 | 17.68 MHz |
mid |
13.5 MHz | 27 MHz | 0 | 19.03 MHz |
fast20 (26-AUG-2026) |
20.25 MHz | 40.5 MHz | 0 | 22.932 MHz (31-AUG; was 20.556) |
full |
27 MHz | 54 MHz | 1667 | 19.55 MHz |
The slow/mid/full Fmax figures above predate the 31-AUG .sdc work
and are therefore PESSIMISTIC by an unmeasured amount - they were taken
while the storage crossings were still being analysed as synchronous. Only
the fast20 row has been re-measured.
31-AUG-2026: fast20 margin went from 1.5% to 13.2%. The CPU-domain
Fmax was never a statement about the CPU. nd120_tang20k.sdc was a single
create_clock line, so the data buses between nd_storage's card side
(clk_stor, the 27 MHz sys_clk domain) and its client side (clk_cpu,
the PLL's CLKOUTD) were analysed as same-edge synchronous paths with a
required time of 0.000 ns. That produced 24 of the 25 worst setup paths
in the build. Two set_false_path lines, measured with the .sdc as the
only variable and the source tree hashed identical either side:
| CPU-domain TNS | failing endpoints | |
|---|---|---|
| without the two lines | -260.076 ns | 398 |
| with them | -6.489 ns | 24 |
Together with aa210eb taking the MIPS tap off ALUCLK_EN (neither fix
alone got there), fast20 now closes at TNS 0.000 on all five clocks,
setup AND hold, CPU Fmax 22.932 MHz against the 20.250 MHz
constraint. Flashed and confirmed running on the board.
Three things learned doing it, all worth knowing before touching the
.sdc again:
- A WIDER exception set measured WORSE. Also excepting the synchroniser
first flops (
*/s_meta*-SD-FAT/circuit/nds_sync.vcalls it// first flop (metastability guard)) and thes_led_*_stretchLED counter describes the hardware just as correctly, but on an identical tree gave -35.352 ns over 90 endpoints vs -6.489/24. Excluding a path also removes the placer's reason to close it. Pick the exception set on measurement, not on how thorough it looks. - Gowin silently ignores a constraint that matches nothing - check
section 3.8
Timing Constraints Reportin the.trforActived. Worse, a constraint naming a register that did not survive synthesis is a hardERROR (TA2003)that aborts the build and leaves the PREVIOUS.tron disk, where it reads as a perfectly plausible result. Compare the.tr<Created Time>against the.sdcmtime. - Do not reach for
set_clock_groups -asynchronousbetween the two domains. It clears every violation in one line and excuses every crossing nobody has read, including the one somebody adds next month.
Only simple ratios of the 27 MHz crystal can be signed off (01-SEP-2026)¶
Measured while looking for a clock above fast20. Four ratios, one build
each, RTL and constraints identical - only the PLL dividers changed:
| CPU clock | ratio to crystal | TNS reported | negative paths in the report |
|---|---|---|---|
| 20.25 MHz | 3/4 | 0.000 / 0 | 0 - agrees |
| 27 MHz | 1/1 | -8.572 / 35 | 35, worst -0.839 - agrees |
| 22.5 MHz | 5/6 | -1332 / 1129 | 0 - contradicts |
| 24.75 MHz | 11/12 | -26678 / 2774 | 27, worst -2.806 - contradicts |
At the two SIMPLE ratios the summary TNS and the path tables agree exactly. At the two awkward ones they contradict each other by two orders of magnitude - 22.5 MHz claims 1129 failing endpoints while its path table holds NONE - and the size of the discrepancy tracks the ratio's complexity (6 edge pairs -> -1332, 12 -> -26678). That is consistent with the tool enumerating launch/capture edge pairs across the common period while report_timing collapses to one representative per endpoint, but the mechanism is INFERRED. What is measured is the correlation.
Constraining every sys_clk <-> CLKOUTD crossing did NOT clear it (24.75 MHz went from -21909/2629 to -26678/2774 with eight more exceptions applied), so it is not the storage crossings.
Consequence: do not ship an intermediate clock. Not because the silicon cannot run it - at 24.75 MHz the worst SAME-domain path was only -0.271 ns - but because the report cannot be trusted to say so, and a bitstream signed off on a number you cannot trust is the thing this whole exercise was correcting. The usable rungs are 6.75, 13.5, 20.25 and 27 MHz.
27 MHz remains 0.839 ns short on 13 endpoints, all WCS -> WCS by the microcode-JUMP and TVEC routes. Those are real; see commit 9dc8507 for why they must not be constrained away.
Console: 115200 baud 7E1 on every variant since 27-AUG-2026
(UART_BAUD_RATE is unconditional in src/tang20k_defines.v:554). The
physical baud is the UART_BAUD_RATE build constant alone; the microcode's BAUDV thumbwheel value (8 = 9600) is stored
by the SC2661 emulation but never used for bit timing, proven on the Nexys
and now here. Silicon 26-AUG-2026: SINTRAN III boots on fast20,
banner + Watchdog in ~40 s, clean text on a 115200 7E1 console. Soaked
27-AUG: 4 unattended hours, 8/8 console probes. All the python console tools
in this directory default to 115200 (--baud 9600 for pre-27-AUG bitstreams).
Measured on silicon, SINTRAN III booting from WD0¶
Cold boot each time (reflash, then 20500&), driven by measure_s3.py.
banner = to SINTRAN III RUNNING; watchdog = to
ERS/SINTRAN III Watchdog has started, i.e. ready for login; S3 = from the
CR that submits S3 to its first output byte, after SET-T-T,,93.
All figures are SECONDS (decimal), not minutes:seconds. 2.4 sec means two
point four seconds - S3 responds almost immediately at every clock. The
minute:second equivalents are given in brackets for the longer ones.
| CPU clock | banner | watchdog (login ready) | S3 first output |
|---|---|---|---|
| 6.75 MHz | 168.2 sec (2 min 48) | 539.3 sec (8 min 59) | 3.9 sec |
| 13.5 MHz | 119.0 sec (1 min 59) | 480.1 sec (8 min 00) | 2.7 sec |
| 27 MHz | 101.9 sec (1 min 42) | 451.9 sec (7 min 32) | 2.4 sec |
Boot is NOT CPU-bound. Four times the clock buys only 1.65x on the banner
and 1.19x on time-to-login. The banner->watchdog segment barely moves at all -
371 s, 361 s, 350 s - so that phase is essentially clock-independent. That is
consistent with it being disc-bound: the SD/storage stack deliberately runs off
the fixed 27 MHz crystal (clk_stor = sys_clk) regardless of the CPU clock,
because sd_file_reader's identification divider is only in spec there.
S3 starts in under 4 seconds at every clock. A "slow S3 start" is therefore not S3 being slow - it is almost certainly the machine still being in the long post-banner phase, before the watchdog line says it is ready. Wait for the watchdog before concluding anything about S3.
Which variant to use¶
full (27 MHz) runs SINTRAN, LIST-FILES and S3, and is the fastest measured -
but it does NOT close timing (1667 setup violations against a 19.55 MHz Fmax),
so its margin over temperature and voltage is unquantified. mid (13.5 MHz)
closes with zero violations and gives most of the gain: 1.41x on the banner
against slow, versus 1.65x for full.
The full row predates both the 31-AUG storage-crossing exceptions and the
01-SEP WCS -> ACAL exception now in the .sdc (its comments give the proof),
so the full figure is a floor, not a verdict. fast20 is the timing-clean
choice.
The debug dumper's baud divisor used to be a fixed DELAY_FRAMES(1406) that
was right only for slow; fixed 24-AUG-2026, it is now derived from
BOARD_CLK_FREQ and UART_BAUD_RATE (src/ND120_TANG20K_TOP.v:1747).
LEDs (active low, pins 15-20; map of 07-AUG-2026)¶
| LED | Meaning |
|---|---|
led[0] |
Storage BLOCK READ - flashes ~150 ms per floppy/Winchester block read (FDISK_REQ/WDISK_REQ with WR low), solid under sustained reads |
led[1] |
Storage BLOCK WRITE - same stretcher, for block writes (WR high) |
led[2] |
A tape byte was served - the SD -> TAPE-400 path delivered data (400$ working) |
led[3] |
sd_status[0] together: 00 = mount never ran, 01 = no card, |
led[4] |
sd_status[1] / 10 = mount/FAT error, 11 = SD-FAT OK (both lit) |
led[5] |
Heartbeat ~0.8 Hz - clk_cpu alive at all |
Reading it: led[4]+led[3] both lit = the whole SD-FAT chain works. After
400$, led[2] dark = the CPU never got tape bytes. led[0]/led[1] are
the disc activity lights: one flash per block, a steady glow during a
transfer burst. (Before 07-AUG, led[0]/led[1] were the bring-up
indicators tape-request-seen / SD-clock-seen; those are retired.)
Storage build: SD-FAT + floppy + SMD (measured 3-AUG-2026)¶
The SD-FAT reader, the floppy at 1560 and the SMD disc at 1540 are all in one bitstream and it places and routes. This is the build to use for disc work.
Defines (src/tang20k_defines.v, both active):
| Define | Effect |
|---|---|
TANG_FLOPPY |
floppy-only base build: ND_FLOPPY_DMA at 1560 + nd_storage client for FLOPPY1.IMG. Drops the papertape - TANG_INC_TAPE goes to 0, so 400$ is not available in this bitstream. |
TANG_SMD |
adds ND_SMD at 1540 with its own ND_DMA_MASTER, plus nd_storage_disc_adapter serving SMD0.IMG (client 3, slot 1376 blocks -> image limit 2,818,048 bytes). |
They resolve to TANG_INC_FLOPPY = 1, TANG_INC_SMD = 1, TANG_INC_TAPE = 0
in src/ND120_TANG20K_TOP.v, applied both to the core (which devices exist) and
to nd_storage_devices (which client the SD-FAT reader serves).
With either define set, the SD-FAT slimming cut SDFAT_NO_STORAGE_CHECK is
suppressed automatically - floppy and SMD do random access and writeback, so the
mount-time contiguity checker (nd_storage_fatchk.v) must stay in. That cut is
for the tape-only build.
Measured utilization (Gowin EDA flow, VARIANT=slow, 3-AUG-2026, from
build/nd120_tang20k_build/impl/pnr/nd120_tang20k_build.rpt.html):
| Resource | Used | % |
|---|---|---|
| LUT/ALU/ROM16 | 14464 (12955 LUT, 1509 ALU) | - |
| CLS | 9103/10368 | 88% |
| Register | 7579/15915 | 48% |
| - as Latch | 0/15552 | 0% |
| - as FF | 7529/15552 | 49% |
| BSRAM | 34 SP10 SDPB | 96% |
| DSP | 2 MULTALU36X18 | 9% |
| PLL | 1/2 | 50% |
Read those two numbers together: BSRAM 96% and CLS 88% is what "everything
fits, with nothing to spare" looks like on this board. The SMD's 1024x16 sector
buffer is the last BSRAM block; dropping TANG_SMD alone is the intended escape
hatch if something else needs one. Register as Latch is 0 - no inferred
latches, which is a standing gate for every build here.
Floppy and SMD fit because both 2 KB sector buffers are synchronous-read (the old asynchronous read ports could not map to BSRAM).
Build and program (Gowin EDA flow, Windows host):
.\gowin_build.ps1 -Variant slow # or: make gowin VARIANT=slow
make load-gowin # SRAM - gone at power-off
make flash-gowin # SPI flash - survives a power cut
Power-cycle the board after every loader operation before judging anything on the console.
Note before flashing rather than loading: this build's SMD write path is live
(full aligned 1024-word blocks only; anything else answers disk_err), and a
flashed bitstream comes up on its own at every power-on, so it can write
SMD0.IMG on the card without anyone asking.
Files here (legacy)¶
| File | Purpose |
|---|---|
ND120_TOP.cst |
STALE - Tang Nano 9K pinout (clock pin 52, LEDs 10-16). Superseded by src/nd120_tang20k.cst. Kept only until the 9K is ever targeted; do not use for the 20K. |
sdram-test/ |
Standalone SDRAM bring-up test - nand2mario controller + ND-120 UART (9600 8N1) reporting every read/write. Gowin EDA project + OSS Makefile + iverilog testbench. PASSES on hardware (2026-07-08, OSS-flow bitstream, full 8 MB write+verify OK). |
sdram18-test/ |
Standalone sdram18.v hardware test - drives the ND-120 18-bit-word controller with the full build's exact 13.5 MHz slow-bring-up clocking (same PLL module). 4-word demo + full 2M-word write/verify over UART. PASSES on hardware (2026-07-09, OSS flow) - exonerates the controller; the deposit bug was full-build cross-domain timing (fixed; the handoff is in git history). |
sdram-bridge/ |
ND-120 sheet-49 SDRAM backend - MEM_RAM_49_SDRAM.v maps the measured ND-120 DRAM protocol onto the SDRAM (2x-clock bridge, self-scheduled refresh, 2 banks = 4 MB). In the booting build (4 MB main memory). Design doc: ../../docs/nd120-dram-memory.md. |
Older Gowin EDA project: ../../ND-120-Gowin/ (ND-120-Gowin.gprj). The
build scripts in this folder (gowin_build.ps1/.tcl, Makefile) use
nd120_tang20k.gprj instead; whether to delete the older project is an open
decision.
Memory architecture (the key design point)¶
BSRAM (828 Kbit) is too small to hold both the microcode PROM (~512 Kbit) and the WCS (~512 Kbit), unlike the Basys3. So:
| Memory | Where on Tang |
|---|---|
| Microcode | Bitstream-preloaded WCS via SKIP_WCS_LOAD (see ../../docs/skip-wcs-load.md); no PROM BRAM (never read once load is skipped). |
| Writable Control Store (WCS) | BSRAM (32 blocks) |
| Main memory | 8 MB SDRAM via the nand2mario Tang-Nano-20K controller, through sdram-bridge/ |
BSRAM is the binding resource on this board - 96% in the storage build
above. The one reclaim still open (repack the UUA half of the WCS, 8 blocks) is
analysed in BSRAM-BUDGET.md, not implemented.
Resource budget (what uses the GW2AR-18, measured 16-JUL-2026, OSS flow)¶
- BSRAM: the WCS is 32 of the 46 blocks (32 x IDT6168A_20 4096x4,
CPU_CS_WCS_21_22.v); MMU page table + cache 8, MMU IMS1403 1, and each 2 KB storage sector buffer 1. Main memory is in SDRAM (0 BSRAM); the CPU register file is LUT RAM. Reclaiming 8 WCS blocks:BSRAM-BUDGET.mdPart 1. - LUT: the SD-FAT/storage stack roughly doubled the LUT count (42% -> 88%
on 16-JUL);
sd_file_readeralone was about 8930 LUTs. By 28-SEP-2026 the OSS synthesis of the full CPU needs 107-109% LUT4 and no longer places (see "Two build flows"); the Gowin flow still fits. The FAT reader is shared by all devices, so another device adapter is cheap (~176 LUTs). - Levers, with measured or estimated savings:
SDFAT_NO_LFN(~1800 LUT, in use), mount-time contiguity checker (~1177 LUT, retired as a default 07-AUG-2026),SDFAT_NO_WRITE(~1081 LUT, loses write-back),N_CLIENTSof nd_storage (LUT/CLS relief), WCS repack (-8 BSRAM, not done). Dropping FAT32 saves ~0 (the 32-bit cluster path serves FAT16 too). Subdirectory traversal does not exist (root only), so there is nothing to cut.
On-chip debug¶
- THE working method:
TRACE-CAPTURE-GUIDE.md- 512-sample on-chip analyzer in the ND-120 top, dumped over the console UART; full build/capture/decode walkthrough. This is what cracked the memory-write bug. - GAO (Gowin Analyzer Oscilloscope) - captures internal nets to BSRAM, read back over JTAG. But BSRAM is scarce here (WCS uses most of it), so GAO capture depth is shallow.
- Preferred: UART debug streamer (BSRAM-free) for full-length boot traces; fast Gowin roundtrip makes re-flashing to move probes cheap.
Related docs¶
TRACE-CAPTURE-GUIDE.md- on-chip trace capture + analysis how-to (this board).../../docs/skip-wcs-load.md- preloaded-WCS microcode (needed for the fit).../../docs/build-defines.md- the board-target define scheme.../../docs/fpga-debug-methodology.md- shared FPGA debug workflow.