Plan: eliminate all 47 unconstrained-clock warnings (17 rogue clock nets)¶
Full path: Verilog/docs/plan-fix-unconstrained-clocks.md
Date: 9-JUL-2026. Solved 10-JUL-2026 (P1-P4 done, 0 x Gowin TA1117,
every real Tang clock closes with 0.000 TNS setup and hold, and the first
console memory WRITE worked on silicon). This file is kept because about 70
RTL files cite its phase names (P1a-P1e, P2, P2b-P2e, P3, P4) for the
conversion each one carries. The phase plans, per-step checklists, the
validation harness and the day-by-day log are in git history; the sim gates
they used are now in make test / make test-full.
Why it was needed: the Tang deposit bug was measured down to cross-clock-domain timing: every ND-120 layer, the SDRAM bridge FSM, and sdram18.v were proven correct in isolation; the full build failed with per-bitstream-stable wrong reads - the signature of unconstrained register-as-clock domains. The Gowin log's 47 "WARN (TA1117) can't calculate clocks' relationship" were 47 symmetric pairs over 17 nets used as clock pins.
Still open (P5): a post-build check in gowin_build.tcl that fails on
any TA1117 (none exists yet), and the equivalent check_timing pass on the
Basys3/Vivado build. Also: the BD-bus strobes (CLKBD/SPEA/SPES, mode 2) are
dead logic on both boards - a residual risk, see P3.
The 17 rogue clock nets¶
| # | Net (Gowin name) | Clocks what | Phase |
|---|---|---|---|
| 1 | s_clk (CYC CLK) |
DGA F924s (incl. the WRITE flop), IO consumers | P2e |
| 2 | s_uclk_Z |
microcode pipeline registers | P2c |
| 3 | s_mclk_Z |
MCLK-domain registers | P2d |
| 4 | s_aluclk_Z |
ALU registers (CGA_ALU_*) | P2b |
| 5 | s_rfclk |
WRF write strobes | P2a |
| 6 | s_clk3_n_10 |
DGA internal inverted clk3 | P2e |
| 7 | MEM/s_rdata |
MEM_DATA_46.v:219,245 AM29861A read latches | P1a |
| 8 | s_refrq_n |
MEM_ADEC_45.v:175 D_FLIPFLOP | P1b |
| 9 | s_dbg_memw_0[2] = ECREQ (v4 dbg layout bit2; NOT MWRITE50_n) |
MEM_ADEC_45.v:235,238 .CK(s_ecreq) |
P1c |
| 10 | BGNT_n_9 |
BIF_DPATH_BDLBD_10.v:75,94,113 TTL_74646/648 .CLKBA(s_bgnt_n) |
P1d |
| 11 | DSTB_n_34 |
BIF_DPATH_CDLBD_11.v:77,93 TTL_74646 .CLKAB(s_dstb_n) |
P3 |
| 12 | BIF/SPES_12 |
BIF_DPATH_PESPEA_13 parity registers | P3 |
| 13 | s_ibapr_n_Z |
BIF bus-address capture | P3 |
| 14 | DELILAH/s_ldirv_2835 |
CGA_CPU_ALU_CONTR.v:657,669 IR registers | P3 |
| 15 | IO/s_sioc_n |
IO_37 registers | P3 |
| 16 | DCD/s_div_16, s_XRTOSC |
IO_DCD_38.v:403 RTC divider chain | P3 |
| 17 | sys_rst_n |
sys_rst_n_s0/F (per .tr; find RTL flop in P4) |
P4 |
P0 inventory addendum (9-JUL-2026, RTL clock-port sweep)¶
Additional register-as-clock consumers found by grepping every clock-capable
port hookup (.CK/.CP/.CLK/.clock/.CLKAB/.CLKBA/.LE/.RAS_n/.CAS_n) in
CPU-BOARD-3202/circuit, DECODE-GateArray, DELILAH-CPU. These do not all
appear as separate TA1117 clocks (some merge or optimize in synthesis) but
they are the same disease and get fixed in the same phase as their neighbors:
| Net | Clocks what | Phase |
|---|---|---|
s_dbapr |
MEM_ADEC_45.v:270,273 .CK(s_dbapr) |
P1c |
s_ddbapr |
MEM_ADEC_45.v:163 .clock(s_ddbapr) D_FLIPFLOP |
P1c |
s_spesl |
MEM_ERROR_47.v:114 .CK(s_spesl) |
P3 |
| RAS_n/CAS_n | MEM_RAM_49.v:170-272 SIP1M9 row/col/write capture | P1e |
s_clk1/s_clk2/s_clk3(_n) |
DECODE_DGA_COMM.v:959-1146 F924 banks (internal CYC-derived) | P2e |
s_mclk as LE |
CPU_PROC_32.v:324 .LE(s_mclk) AM29841 latch |
P2d |
Clock Summary cross-check (build/nd120_tang20k_build/impl/pnr/*.tr 2.2):
19 base clocks = sys_clk + the 18 rogue roots; matches the table above plus
s_ldirv (source DCD/GATES_10/s_ldirv_s3/F - a LUT output as clock).
Max-frequency reality check from the same report: s_clk domain Fmax
2.689 MHz, s_mclk 2.732 MHz, s_aluclk 5.286 MHz - even the
constrainable paths are 5-10x too slow for the 13.5 MHz clk2x, and TNS
on aluclk is -40862 ns over 355 endpoints. P2 does not just constrain
these paths, it moves them onto sysclk where the tools can finally retime.
Rules that bind every phase:
- Never modify Verilog/PAL/PAL_*.v - conversions live in consumers or
_D wrapper modules (CYC_CC_D / PAL_44446B_D pattern).
- Convert a clock domain whole, never partially (R41P lesson: a
half-converted domain creates phase skew between its own registers).
- One net / one domain per commit - every step revertable and bisectable.
- All conversions sit behind FPGA_FF_MODE; latch mode (original hardware
semantics) stays byte-identical.
Phase 1 - memory-path capture points (deposit-critical, smallest blast radius)¶
Converted #7-#10 to the proven sysclk edge-capture pattern (AM29C821
USE_SYSCLK=2 - one capture per detected CK rise, no routed clock; the same
fix that solved the MEM_ADDR_44 deposit regression):
- P1a MEM_DATA_46 RDATA read latches (add USE_SYSCLK-style mode to AM29861A, matching AM29C821's).
- P1b MEM_ADEC_45:175 refrq_n flop -> sysclk + refrq edge detect.
- P1c MWRITE50_n-as-clock consumer (from P0 inventory).
- P1d BGNT_n-as-clock consumer (from P0 inventory).
- P1e SIP1M9 / MEM_RAM_49 RAS-edge row capture -> sysclk + RAS edge detect (Basys3 evidence: PAL_44902 RAS_n is an unconstrained root clock feeding the BRAM row capture - deposit-critical on the Basys3; see the Basys3 appendix).
Phase 2 - the CYC_36 generated CPU clocks (the branch thesis)¶
CYC_36 additionally emits one-sysclk-wide enable pulses aligned with each
level-registered clock: CLK_EN, UCLK_EN, MCLK_EN, MACLK_EN, ALUCLK_EN
(landed 9-JUL - see progress log; alignment property tb = test-cycen).
Consumers convert from posedge s_xclk to posedge sysclk + if (XCLK_EN),
one whole domain per commit, smallest first.
CORRECTION (9-JUL, from the P1d .tr): the plan's original "P2a s_rfclk
(WRF)" was a misidentification - netlist s_rfclk's source is
DGA/POW/A633/reqQ_n (the DGA power/RTC divider chain, 1 endpoint), NOT
the CGA write register file. It moves to the P3 IO/divider batch. The WRF
is inside the ALUCLK/WRFSTB structure and is handled with P2b.
- P2b s_aluclk domain (CGA_ALU: GPR/DBR/QREG/STS + CGA condition regs)
- P2c s_uclk domain (microcode pipeline)
- P2d s_mclk domain
- P2e s_clk domain (DGA XCLK + IO). Converting the DGA turns F924 into sysclk+CE and dissolves the internal clk2/clk3/clk3_n (#6) for free.
REVISION (10-JUL, from the P2b red gates): P2b-P2d cannot land as separate commits - the CGA clock group must convert TOGETHER. Proof by measurement:
- With rise-aligned enables the P2b conversion is byte-identical to golden
for 557,335 ticks, then ONE unconverted flop flips:
CGA_MIC.MEMORY_35(s_zfff_q_out, Z-flag sync,posedge s_mclk, d = combinational ZF off the converted ALU registers). Waveforms show MCLK and ALUCLK rising on the SAME sysclk edge in that cycle shape: the original derived-clock flop samples ZF BEFORE the ALU registers' same-edge update lands (NBA old-value exchange between two pass-2 derived clocks); the converted ALU registers update in the sysclk pass, so the still-pa-clocked MEMORY_35 captures the NEW ZF one clock early. The corruption spreads ZFFF -> COND -> INTR priority state (PD/HIGSN at ~650k) -> self-test TEST 2 STERR at ~737836 -> runSim WAIT-loop hang. - Enable placement cannot fix this class. All three placements were built and measured: rise-aligned (correct for converted<->converted and latch-fed data; breaks unconverted same-edge observers), rise+1 posedge and rise+negedge (both shift the capture VALUE to post-edge, which diverges from cycle 0 - BDRY refresh chain +1 tick, 900k-line diff). Verilator's derived-clock flops sample PRE-edge values of other same-edge derived-clock flops; only same-pass NBA reproduces that, i.e. BOTH sides of every coincident-edge transfer must be converted.
- ALUCLK/CLK rise coincide by construction (both TERM-derived) and MCLK
coincides in short cycle shapes, so the coupled unit is the whole CGA
clock group: ALUCLK + CLK + UCLK + MCLK + MACLK consumers (MIC, INTR,
TRAP, IDBCTL, DCD, MAC, ALU, WRF) in ONE tree state, gated by the same
byte-identity compare at the end. Fall-edge consumers
(InvertClockEnable / ~ALUCLK style) need FALL enable pulses
(
pa & ~next) added to CYC_36 alongside the rise enables. - Debug harness for residual flips (sim/):
ND120_START_TRACE/ND120_MAX_TICKSenv overrides in test_nd120.cpp + kept obj_dir_base (conversions off) and obj_dir_conv builds + scratchpad vcd_diff.py/vcd_snap.py/bisect_time.sh - binary-searches the first diverging tick and names the first diverging signals in ~30 min without rebuilds. Ground zero for any new flip = convert that observer's module next.
GROUP CONVERSION (10-JUL): the whole coupled group is now converted with rise-aligned enables:
- CGA: ALU + WRF (P2b) + MIC (48 flops incl. M169C LC counters, SR44 stack on MCLK-fall) + INTR (17) + DCD (21) + IDBCTL PGSREG (14) + MAC (7+1) + TRAP TVGEN_P2 (7, TCLK=UCLK domain). New wrappers: J_K_FLIPFLOP_EN, SR44_EN, SCAN_WITH_SET_N_EN, SCAN_WITH_RESET_N_EN, M169C_EN, F924_EN (composites = structural copies with only the flop primitive swapped).
- CYC PALs 44403C/44404C: new PAL_44403C_EN.v / PAL_44404C_EN.v in
Verilog/PAL/ (equation copies with
if (EN); originals untouched) - they register on CLK and sample converted outputs (ACOND_n -> LCS_n was the 557151 detonation). - DGA (P2e pulled in): COMM (13) + IDBS (4) + panel FIFO on XCLK=CLK; one CLK_FALL_EN site (COMM A204 on ~clk3). POW divider left for P3.
- Board: ND3202D regMIS, IO_UART_42 CHIP_33G (AM29C821 CK driven by the CLK_EN pulse in FF mode). CS/ACAL/PANCAL verified no pa-clocked flops.
- CYC_36 emits FALL enables (level & ~next) for all five clocks; test-cycen extended (18990 checks).
- Wrapper init bug found by harness: SCAN_WITH_SET_N_EN async set must be
posedge ~S_n(active-high preset like the original ACTIVE_ASYNC flop), NOTnegedge S_n- at time 0 S_n starts low, the negedge never fires, the original's preset 0->1 does (DZD_FF t=0 flip). - BDRY baseline finding: latch-vs-FF
make compareat committed HEAD (0843e88) already differs in EXACTLY the BDRY column (5872 rows, refresh pulses at reset rows 0-7 in latch, 100-107 in FF, then +1 offset) - a PRE-EXISTING FF-mode reset artifact, not from this campaign. The practical gate is therefore: diff columns == {BDRY, first@0, 5872} and nothing else. - With all of the above, latch-vs-FF divergence moved 557151 -> 738965 (only CSA/LED/MCLK/TERM_n after 738965 + the baseline BDRY artifact).
ALL SIM GATES GREEN (10-JUL): after two more fixes the full group conversion passes everything:
- CMDDEC PALs 44407A/44408B/44511A converted (PAL_*_EN equation copies, CLK_EN threaded ND3202D -> CPU_15 -> CPU_PROC_32 -> CMDDEC). Fixed the VEX/LDEXM flip at ~557533.
- CGA_MIC_MASEL regIW: its FF-mode capture read regREP - but regREP is itself a sysclk register updating on the same edge, so the pa-clocked original saw regREP's NEW value while the converted pre-edge sample was one cycle stale. Fix: capture regREP_comb (the register's D input). This is THE pattern for native-sysclk-producer -> converted-consumer same-edge transfers: sample the producer's D, not its Q. Fired only when a conditional jump resolved exactly on an MCLK rise (survived to tick ~738920, self-test).
- Result: latch and FF traces byte-identical to the STORED GOLDENS
(cmp trace_{ff,latch}.csv vs sim/golden/ = OK over 1M rows), global
make test24/24, runSim FF console byte-identical to golden (previously hung in the STERR wait-loop), Tang vtest deposit 22/054321 readback PASS.
Fmax reality from the P1d .tr (why P2 matters): s_clk 3.06 MHz, s_mclk_Z 3.07 MHz, s_aluclk_Z 5.85 MHz, CLKOUTD domain 5.67 MHz vs the 6.75 MHz crawl clock - and none of it hold-analyzed.
Phase 3 - control-strobe clocks (#11-#16) - done 10-JUL-2026¶
Same edge-capture conversion, batched per module: BIF batch (DSTB_n, SPES, ibapr_n), CGA batch (LDIRV), IO batch (sioc_n, div_16, XRTOSC divider chain -> sysclk counter with terminal-count enables). TA1117 went 28 -> 2.
- BIF end-of-window lesson (the board deposit bug): DSTB_n marks the
END of the memory data window - the mode-2 (rise+1) edge capture reads
a dead bus, and every memory READ returned 000000 (vtest + runSim both
caught it; examine looked right only because memory was zero).
New
TTL_74646 USE_SYSCLK_AB=3WINDOWED capture: follow the data on posedge sysclk while the strobe is low, hold from the rise -> holds the at-rise value with no lag. Legal in CDLBD because SAB=DSTB_n selects real-time data during the low window, so regA is unobserved until the rise. RULE: strobe conversions must classify the strobe first - START-of-window (ECREQ, BGNT_n: data valid after the rise -> mode 2) vs END-of-window (DSTB_n: data dies at the rise -> mode 3). - BIF batch: CDLBD CLKAB (mode 3), BDLBD CLKAB=CLKBD (mode 2, BD bus tied off - residual risk documented), PESPEA 4x TTL_74534 on SPEA/SPES (new USE_SYSCLK=2 param), PPNLBD posedge ECREQ -> inline edge capture. s_dbg_memw_0[2] root was ECREQ seen through the DBG_MEMW bus - gone with the PPNLBD fix.
- IOC: TTL_74273 gained USE_SYSCLK=2 (sync clear kept); CHIP_28A_IOC captures on the detected ~SIOC_n rise. Cost: the two LED trace columns shift one tick (the only non-BDRY trace change - accepted, FF golden regenerated).
- LDIRV: D_FLIPFLOP_EN gained USE_ENABLE=2 (strobe edge-capture via the clock pin); MEMORY_46/47 in CGA_CPU_ALU_CONTR converted.
- Dividers: the two ripple 74393s (s_div_16/s_XRTOSC roots) replaced in
FF mode by one synchronous 8-bit counter (bit-exact mapping: pposc=
bit2, div_16=bit3, XRTOSC=bit7); the POW F714/F617 RTOSC ripple network
(s_rfclk root) replaced by a sync 6-bit counter + rfclk/panosc toggles
- F617 set-priority equations, with the rfclk rise derived from the q633 NEXT value so the CLOSC-forced QB rise still clocks A630/A631. Refresh phase moves -> BDRY diff rows grew (13684, still BDRY-only).
- MMU cache roots (s_mclk_Z/s_uclk_Z): PAL_44402D posedge UCLK -> PAL_44402D_EN equation copy on UCLK_EN; AM29841 CHIP_25F posedge LE=mclk -> new AM29841 USE_ENABLE=1 on MCLK_EN. UCLK_EN threaded CPU_15 -> CPU_MMU_24 -> CPU_MMU_CACHE_25. Cache RAMs were already sysclk BRAMs.
- vtest caught a P3 regression before the board did: mode-2 DSTB capture read a dead bus and ALL memory reads returned 000000 -> the mode-3 windowed capture above (TTL_74646 USE_SYSCLK_AB=3) fixed it; both goldens stayed byte-identical.
Phase 4 - sys_rst_n as clock (#17) - done 10-JUL-2026¶
- After the P3 fixes the board stayed silent: hardware bisect (DIAG builds) showed the boot lived or died with the SIOC conversion - but DIAG2's timing report was violation-free and the identical RTL passes vtest. The real mechanism was the LAST unconstrained root: POW A572 clocked by s_clear_n (= sys_rst_n on FPGA). Gowin auto-created a bogus 100MHz "sys_rst_n" base clock for it and every path near it was unanalyzed - the same per-bitstream placement-lottery that caused the original write bug. Each recompile rolled the dice; the SIOC localparam flip just reshuffled the layout.
- P4 fix: A572 converted in FF mode to a sysclk flop capturing s_esload_n on a detected s_clear_n rise (async CLRTI preset kept, q init 0 so s_lrst starts 1 like the original).
- RESULT: zero TA1117 warnings, the bogus sys_rst_n clock is gone from the Clock Summary, and every real clock closes timing with 0.000 TNS setup AND hold. All sim gates green (24/24 units, both trace goldens byte-identical, runSim console golden, vtest PASS). BOARD: boots to the OPCOM '#' prompt, deposit 22/54321 + readback = 054321 - the first time a memory WRITE from the console works on silicon.
Also found on the way: the write-path analyzer in ND120_TANG20K_TOP
triggered on the first write decode - which fires during a normal boot once
writes work - and then held the console TX pin forever. That pin takeover
is now behind TANG_WRITE_ANALYZER_DUMP (off by default).
Gowin SDC lessons (from the thrown-away P0.3 constraint probe)¶
The probe (constrain the derived clocks instead of converting them) did not fix the board: a netlist whose real Fmax is below its clock cannot be rescued by constraints. The Gowin SDC lessons it cost, three builds' worth:
1. The SDC parser is a tiny Tcl subset: no foreach/variables -
flat commands only.
2. PLL output pins are not clock sources until explicitly declared -
chain create_generated_clock from [get_ports sys_clk] first.
3. Auto-inferred clocks are not addressable objects - get_clocks on
a .tr auto-clock name fails with TA2004 at parse time; declare an
explicit create_clock on the source PIN (from the .tr Source column),
then reference the explicit name.
4. Synthesis uniquification suffixes in pin names change between builds
(s_dbg_memw_2_s17/F in one build, _s2/F in the next), so a
pin-named constraint is fragile.