ND-120 on Nexys 4 DDR - Timing Closure Report (clock-up campaign)¶
Historical / campaign closed (kept as the measured timing record). The STA numbers below stand. Since this was written, SINTRAN boots on silicon at 45.45 MHz and 50 MHz, and the board is deployed at 33.333 MHz with the cache ON (the cache added routing pressure that reopened 45.45 - see
../timing.md, sections dated 30/31-AUG-2026). Read this report for the per-domain analysis and the frequency search, not for the "next steps".
Started 26-AUG-2026, branch clock-up. All paths in this document are
relative to the repo root unless stated otherwise. Evidence directories:
Verilog/fpga/nexys4ddr/timing-analysis/baseline_clk16/ (preserved reports
of the shipped build) and timing-analysis/run_clk<N>/ (one per
implementation run, generated by the battery in build.tcl).
1. Executive summary¶
The shipped SINTRAN-booting build runs the CPU at 16.667 MHz with the CPU domain only ~55% utilized in time: its worst path arrives in 33 ns of a 60 ns budget. A ten-run post-route frequency search (every run a full synth->place->route with the real constraint, never synthesis arithmetic) measured the actual wall:
- Every step from 25.0 to 45.45 MHz closes timing with the default flow. 45.45 MHz (22 ns) closes at WNS +0.085 ns; 50 MHz (20 ns) fails the default flow by -2.546 ns across 1213 endpoints.
- Adding
phys_opt_design(new opt-in-tclargs physopt) closes 50 MHz at WNS +0.007 ns - 3x the shipped clock, zero practical margin. - The CPU domain has exactly ONE critical-path family at every frequency: the full microcycle from the WCS microcode BRAM through IDB decode/ALU/TRAP/PALs to the register-file clock-enables, 29-30 logic levels, routing-dominated (76% at 60 ns, 64% at 22 ns).
- One stale constraint was found and fixed (crossing bounds into the CPU clock hard-coded at 80 ns; now track the selected period). No timing exception was added anywhere.
- STA pass is not functional proof. (When this report was written nothing above 16.667 MHz had booted; since then 45.45 MHz and 50 MHz have booted on silicon, and the deployed build is 33.333 MHz with the cache ON.) The two auto-inserted loop-breaking false paths (CGA IDB ring remnants) make every CPU WNS a floor; and WNS +0.007 (50 MHz) or +0.085 (45 MHz) leaves no allowance for the untimed paths, multi-run spread, or anything else. The defensible recommendation is below (section 9).
Recommended: 25 MHz as the new proven-margin default candidate (WNS +9.3 ns, 23% margin), 33.3 MHz as the balanced target (WNS +1.28 ns, 2x the shipped clock) - each contingent on booting SINTRAN and running the instruction-verify suite on the board at that clock, which requires flashing (Ronny's call).
2. Exact board / FPGA / tool identification¶
Everything below is read from the project files and the routed report, not inferred:
| Item | Value | Evidence |
|---|---|---|
| Board | Digilent Nexys 4 DDR (Rev C) | Verilog/fpga/nexys4ddr/nd120_nexys4ddr.xdc header ("Pin source of truth: Nexys-4-DDR-Master.xdc, Digilent official, Rev. C board"); DDR2 MIG core ddr2-test/ip/ddr/ddr.xci (the non-DDR Nexys 4 has cellular RAM, no DDR2) |
| FPGA part | xc7a100tcsg324-1 (Artix-7, speed grade -1) |
build.tcl line 19 (set part); routed report header "Device: 7a100t-csg324, Speed File: -1 PRODUCTION 1.23 2018-06-13" |
| Vivado | v2026.1 (win64, Build 6511674), Windows host | timing-analysis/baseline_clk16/timing.rpt header |
| Flow | Non-project in-memory batch (create_project -in_memory), synth -> opt -> place -> route in one script |
build.tcl |
| Top module | nd120_nexys4ddr_top |
build.tcl synth_design -top; Verilog/fpga/nexys4ddr/nd120_nexys4ddr_top.v |
| Oscillator | 100 MHz on pin E3 (clk100) |
nd120_nexys4ddr.xdc: create_clock -name sys_clk -period 10.000 |
| Baseline report state | Routed ("Design State: Routed"), multi-corner (Slow + Fast), pessimism removal on | baseline_clk16/timing.rpt lines 1-35 |
| Report vs sources | The 25-AUG 21:05 report matches the committed sources: the last RTL/constraint commit precedes the 21:02 build, and later commits (ebe018d, e20678d, 46b5e32) are documentation/submodule only |
git log |
| Build config of the baseline | clk 16 ilaslim : 16.667 MHz CPU clock, DDR2 main RAM (MAIN_RAM_DDR2), CPU cache off (ND120_NO_CACHE), WCS preloaded (SKIP_WCS_LOAD), FF mode, slim ILA + dbg_hub in the netlist |
build.tcl defaults + the shipped nd120_nexys4ddr.ltx existing |
3. Clock architecture¶
One MMCME2_BASE in nd120_nexys4ddr_top.v (line ~108), VCO = 100 MHz x 10
= 1000 MHz:
| Clock | Source | Period / freq | Moves with clk=? |
Drives |
|---|---|---|---|---|
sys_clk |
pin E3 | 10 ns / 100 MHz | no | MMCM input, 7-seg + LED stretchers, heartbeat |
clk_cpu_pre -> BUFG clk_cpu |
MMCM CLKOUT0, CLKOUT0_DIVIDE_F = ND120_N4DDR_MMCM_DIV |
60 ns / 16.667 MHz at clk=16 |
YES - the only one | entire CPU + board + ND-BUS device chain + BRAM-cache side of main RAM |
clk_stor_pre -> clk_stor |
MMCM CLKOUT1, /37 | 37 ns / 27.027 MHz | no | SD/FAT storage stack |
clk200_pre -> clk200 |
MMCM CLKOUT2, /5 | 5 ns / 200 MHz | no | DDR2 MIG reference |
clk_pll_i (ui_clk) |
inside MIG, from clk200 | 13.333 ns / 75 MHz | no | DDR2 user interface, nd_ddr2_port/arb/cache backend |
MIG internal (freq_refclk, mem_refclk, oserdes*, iserdes*, sync_pulse) |
MIG PLL | 300/600 MHz etc. | no | DDR2 PHY |
| dbg_hub TCK | JTAG | 33 ns | no | ILA/dbg_hub only |
The CPU period in ns equals the divider value (VCO 1000 MHz), so the
clk_table in build.tcl maps clk=16 -> 60 ns, 20 -> 50, 25 -> 40,
27 -> 37, 33 -> 30. BOARD_CLK_FREQ moves with the divider so UART
baud, RTC tick and watchdog counts stay correct - the two must move
together (build.tcl comment, verified in the define list).
4. Baseline post-route timing (clk=16, shipped SINTRAN-booting build)¶
From baseline_clk16/timing.rpt (routed, 25-AUG-2026 21:05):
| Metric | Value |
|---|---|
| WNS (global) | +1.460 ns - held by a MIG-internal path sync_pulse -> mem_refclk (report line 224), NOT by the CPU domain |
| TNS | 0.000, 0 failing endpoints of 53987 |
| WHS | +0.024 ns (floppy DMA dma_addr_reg -> s_addr_reg, intra-CPU; hold is period-independent) |
| WPWS | +0.264 ns (clk200 MIG domain) |
| clk_cpu_pre intra-domain WNS | +27.014 ns of a 60 ns requirement |
| clk_stor_pre intra-domain WNS | +13.489 ns of 37 ns |
| clk_pll_i intra-domain WNS | +5.993 ns of 13.333 ns |
| Recovery (async CLR deassert, "async_default" cpu->cpu) | +48.541 ns - reset-recovery checks, healthy |
Worst CPU-domain path (slack +27.014, i.e. 32.986 ns arrival):
CPU/CS/WCS/CHIP_21C RAMB18 (microinstruction word) -> CSIDBS decode ->
INTR/CNTLR/IRGEL -> ALU/OUTMUX + STS + RALU carry -> FIDBO ->
TRAP/BRKDET -> PAL_44307_UCYCLK MAP_n -> MIC/MIC_IPOS MA ->
CS/ACAL LUA -> PAL_44403_UCYIN0 -> TERM_D -> ALUCLK_EN (fo=238) ->
WRF/RBLOCK/Z_REG_0 clock-enable. 30 logic levels, 7.7 ns logic /
24.7 ns routing (76%) - a routing-dominated path placed under a loose
60 ns budget. This is the machine's full microcycle: microcode word out of
the WCS to the write-enable of the register file.
Estimated CPU-domain boundary from this single number: 60 - 27.014 = 32.99 ns -> ~30.3 MHz. This is a timing-model estimate at THIS placement; with 76% of the path being routing under no pressure, a tighter constraint may do better - or uncover harder paths. Only reruns at candidate constraints decide (section 8).
Known caveat that makes every CPU WNS a floor, not a guarantee: synthesis
auto-inserts two loop-breaking false paths through the historical CGA IDB
ring (ALU_OUTMUX/D_15_0[8] and ALU_i_426/O, Synth 8-326;
check_timing reports 6 combinational loops). Paths through those nodes
are untimed. Inherited, documented in build.tcl ("PARKED DEBT"); the RTL
ring cut is the real fix and is out of scope for this campaign.
5. Constraint audit¶
| # | Finding | Verdict / action |
|---|---|---|
| 1 | create_clock on clk100 (sys_clk, 10 ns) present; all MMCM/MIG generated clocks auto-derived (check_timing generated_clocks: 0) |
OK |
| 2 | set_max_delay -datapath_only bounds INTO clk_cpu were hard-coded 80.000 ns (one destination period of the 12.5 MHz era) while the actual period is 60 ns and shrinking. Contract in the comment: "no payload bit may take longer than one destination period" |
FIXED 26-AUG in build.tcl: bound = $mmcm_div (the CPU period in ns), tracks clk=. Safe direction - strictly tighter |
| 3 | nd120_timing.xdc set_clock_groups -asynchronous (sys_clk vs clk_cpu_pre) OUTRANKS the later set_max_delay cpu->sys 10 ns in build.tcl, so cpu->sys crossings are untimed (no cpu_pre->sys row in the Inter Clock Table) |
Accepted: those crossings are single-bit flags through 2-FF synchronizers (nd120_nexys4ddr_top.v sync_hold/sync_grb), where max-delay is a nicety, not a correctness need. Documented, not changed |
| 4 | check_timing: 5 no_input_delay (UART RX, SD card inputs, buttons), 40 no_output_delay (LEDs, 7-seg, UART TX, SD) |
Accepted: asynchronous human/serial interfaces; SD bus is source-synchronous to the FPGA-driven sd_clk in the fixed 27 MHz storage domain, unchanged by this campaign |
| 5 | 6 combinational loops + 2 auto false paths (CGA IDB ring remnants) | Inherited parked debt, see section 4 caveat |
| 6 | Reset recovery/removal: timed (async_default recovery checks), worst arrival 10.6 ns | OK at every candidate period |
| 7 | No multicycle paths declared anywhere | Correct as-is: no path has been PROVEN multicycle by RTL analysis. None added (rule: no exception without functional proof) |
| 8 | MIG constraints come with the core (read_ip), pins included |
OK, untouched |
6. Critical-path families¶
Per-domain extraction (report_cpu_paths.tcl on the run checkpoints,
setup_paths_cpu_group.rpt in each run directory):
The CPU domain has exactly ONE critical-path family. At clk=16 AND at clk=33, all 100 worst CPU-group paths are the same shape:
- Startpoint: WCS microcode BRAM read clock (
CORE/CPU_BOARD/CPU/CS/WCS/CHIP_21C..22D/idt_memory_array_reg/CLKARDCLK, RAMB18E1) - Endpoint: register-file clock-enables (
CORE/CPU_BOARD/CPU/PROC/CGA/DELILAH/WRF/RBLOCK/<R0-R7,Z>_REG_*/regFF_reg[*]/CE, FDRE) - Route (from the clk=16 worst path,
baseline_clk16/timing.rpt): microcode word ->CSIDBS_4_0IDB-source decode ->INTR/CNTLR/IRGEL->ALU/ALU_OUTMUX + ALU_STS->ALU_RALUCARRY4 chain ->FIDBO->TRAP/TVGEN + BRKDET->PAL_44307_UCYCLK(BRK_n/MAP_n) ->MIC/MIC_IPOS(MA_12_0) ->CS/ACAL(LUA_12_0) ->PAL_44403_UCYIN0->U_TERM_D->ALUCLK_EN(fanout 238) -> WRF CE. Functionally: this is the machine's full microcycle - the microinstruction coming out of the WCS deciding, combinationally, whether the register file clocks at the end of the same cycle. - Composition: 29-30 logic levels (LUT2..LUT6, 3x CARRY4), logic delay is
only 7.0-7.8 ns; ROUTING dominates at every constraint tried and keeps
compressing under pressure: 25.3 ns route (76%) at the 60 ns budget,
21.1 ns (75%) at 30 ns, 13.877 ns (64%) at 22 ns (measured:
run_clk45/timing_summary_post_route.rpt, worst CPU path 21.681 ns = 7.804 logic + 13.877 route). - Root cause class: excessive combinational depth (a faithful transcription
of the original board's asynchronous microcycle logic - PALs and TTL
between two clocked elements), amplified by high-fanout control
distribution (
ALUCLK_ENfo=238,CSIDBSfo=35,MA_12_0fo=39).
No second family surfaced anywhere in the search - when the microcycle path
is met, the whole CPU domain is met. The fixed-frequency domains
(clk_pll_i 75 MHz DDR2 user logic, clk_stor_pre 27 MHz SD stack, MIG
PHY) hold the global WNS once the CPU domain is fast, and none of them moves
with clk=.
Supporting observations from the clk=16 battery
(run_clk16/):
qor_assessment.rpt: design "easily meets timing" at 60 ns; no ML suggestions generated.high_fanout_nets.rpt: the largest fanout nets are reset distribution in the storage/CPU domains (2628, 1416, 778, 602) - all with >24 ns slack at their periods; not limiting.cdc.rpt: CDC-8 x2 (reset synchronizers missing ASYNC_REG), CDC-5 x1 (multi-bit synchronizer missing ASYNC_REG) - hygiene items, listed in section 7.methodology.rpt: TIMING-24 x1 confirms one datapath-only bound is overridden (the cpu->sys bound under the clock group, audit item 3); SYNTH-6 x109 (BRAM output timing sub-optimal - the WCS BRAMs read combinationally into the microcycle, no output register: that IS the critical-path family); LUTAR-1 x6 (LUT drives async reset).
7. Improvement options¶
The frequency search (section 8) showed the as-is RTL already closes 45.45 MHz - 2.7x the shipped clock - so no RTL change is NEEDED to raise the clock substantially. Options if more speed or more margin is wanted, ranked:
Low risk (no functional change):
phys_opt_designafter place and after route (-tclargs physopt, added 26-AUG as an opt-in flag). Expected: recovers a few hundred ps on routing-dominated paths; tested at clk=50 (run_clk50_1, section 8). Effort: none (done). Preserves cycle accuracy: yes.- ASYNC_REG hygiene: add
(* ASYNC_REG = "true" *)to the three flagged synchronizer chains (CDC-5/CDC-8 inrun_clk16/cdc.rpt, incl. the debug-panelsync_hold/sync_grbinnd120_nexys4ddr_top.v). Benefit: keeps synchronizer pairs placed together (MTBF), removes the warnings; no Fmax change. Effort: minutes. Risk: none. - Directive/seed sweep near the wall (
place_design -directive,route_design -directive, e.g. ExtraTimingOpt / AggressiveExplore). Vivado is deterministic per configuration, so the single-seed +0.085 at 22 ns says nothing about spread; a sweep establishes "consistently achievable" vs "lucky". Effort: ~40 min per run, no code change.
Medium risk (RTL, changes netlist but not architecture):
- Register the WCS BRAM output (RAMB18E1 DOA register,
DOA_REG=1). Removes the BRAM clock-to-out (2.45 ns) from the family head and lets the BRAM place freely - BUT the microinstruction then arrives one cycle late, which changes the microengine's fetch/execute overlap. On this design that is an ARCHITECTURAL change in disguise: NOT recommended without a full golden-trace revalidation (make compare, instruction-verify campaign).
Architectural (documented only - explicitly out of scope for a cycle-faithful reconstruction):
- Pipeline the microcycle (split at FIDBO or at the MIC address formation). Would roughly double Fmax and roughly halve instructions-per-clock; destroys cycle accuracy against the original machine and every validated timing assumption (RTC, UART pacing, device handshakes). Not recommended for this project's goal.
- Break the remaining CGA IDB combinational ring in RTL (the "PARKED
DEBT" in
build.tcl): removes the 2 auto false paths and the 6 check_timing loops, making the WNS a real guarantee instead of a floor. This is a correctness-of-analysis improvement, not a speed improvement, and the right long-term move (tracked in the 20-AUG analysis,../BUILD-WARNINGS-ANALYSIS.md).
Constraint work (done during this campaign):
- FIXED:
set_max_delay -datapath_onlybounds intoclk_cpunow track the CPU period ($mmcm_div) instead of a stale 80 ns (section 5, item 2). - NOT done, documented: no multicycle exceptions were added anywhere. The 52 ns WCS->ACAL "multicycle" claim from the Tang notes was NOT assumed true here - no RTL proof exists that the path is multicycle, and the search closed 45 MHz without it.
8. Experiments¶
All runs: build.tcl flow, -tclargs clk <N> ilaslim -noburn, full
post-route battery + checkpoint in timing-analysis/run_clk<N>/. Same
config as the shipped SINTRAN-booting build except the CPU divider.
| Run | Change vs previous | clk_cpu WNS | global WNS | TNS | WHS | Result |
|---|---|---|---|---|---|---|
| baseline (25-AUG) | shipped build, 80 ns stale ->cpu bounds | +27.014 (60 ns req) | +1.460 | 0 | +0.024 | reference; SINTRAN boots 7/7 |
| run_clk16 | ->cpu crossing bounds now track the period (60 ns) | +26.455 (60 ns req) | +1.460 | 0 | +0.012 | PASS - tightened bounds cost nothing; CPU worst arrival 33.0 ns |
| run_clk25 | CPU divider 40 -> 25.0 MHz | +9.293 (40 ns req) | +1.224 | 0 | +0.032 | PASS - CPU worst arrival 30.7 ns; global WNS now in fixed 75 MHz clk_pll_i domain |
| run_clk33 | CPU divider 30 -> 33.333 MHz | +1.282 (30 ns req) | +1.282 | 0 | +0.019 | PASS - CPU worst arrival 28.7 ns; CPU domain is now the design-wide binding domain |
| run_clk35 | new clk_table entry: divider 28.0 -> 35.714 MHz | +1.308 (28 ns req) | +1.308 | 0 | +0.023 | PASS - CPU worst arrival 26.7 ns |
| run_clk38 | new clk_table entry: divider 26.0 -> 38.462 MHz | +0.316 (26 ns req) | +0.316 | 0 | +0.013 | PASS by a hair - CPU worst arrival 25.7 ns; margin collapsed, wall is here |
| run_clk40 | new clk_table entry: divider 25.0 -> 40.0 MHz | +0.319 (25 ns req) | +0.319 | 0 | +0.015 | PASS - CPU worst arrival 24.7 ns; router still tracking the constraint |
| run_clk42 | new clk_table entry: divider 24.0 -> 41.667 MHz | +0.152 (24 ns req) | +0.152 | 0 | +0.025 | PASS - CPU worst arrival 23.85 ns |
| run_clk45 | new clk_table entry: divider 22.0 -> 45.455 MHz | +0.085 (22 ns req) | +0.085 | 0 | +0.023 | PASS - CPU worst arrival 21.9 ns |
| run_clk50 | existing entry: divider 20.0 -> 50.0 MHz | -2.546 (20 ns req) | -2.546 | -1600.263 / 1213 endpoints | +0.020 | FAIL - gate refused the bitstream; the wall is between 20 and 22 ns |
| run_clk50_1 | + physopt (post-place & post-route phys_opt_design -directive AggressiveExplore, new opt-in flag) |
+0.007 (20 ns req) | +0.007 | 0 | +0.020 | PASS - phys_opt recovered the full 2.55 ns; 50 MHz is STA-legal with zero practical margin. SILICON 26-AUG: SINTRAN III BOOTS at 50 MHz (this bitstream, one boot, banner reached, console driven by Ronny) |
| run_clk50_2 | console baud 9600 -> 115200 (UART divider constant 5208 -> 434), same clk 50 + physopt | -0.210 (20 ns req) | -0.210 | -4.586 / 54 endpoints | +0.052 | FAIL - the one-constant change re-rolled placement and the +0.007 closure did not survive. 50 MHz closure is FRAGILE: any netlist edit is a new seed |
| run_clk45_1 | clk 45 + physopt + 115200 baud | +0.020 (22 ns req) | +0.020 | 0 | +0.031 | PASS - programmed 26-AUG 10:17, SINTRAN BOOTS on silicon with the 115200 console (Ronny-verified). Deployed configuration |
All runs one seed each (Vivado is deterministic per configuration - spread across directives/seeds is NOT known; see section 9). WHS stayed in +0.012..+0.043 across all runs (hold is period-independent and healthy). WPWS +0.264 (MIG clk200 domain) at every run.
9. Frequency roadmap and remaining work¶
Estimated boundaries per domain (post-route, per-clock WNS, NOT one global
number): the CPU domain is the only one that moves with clk=; the fixed
domains (clk_pll_i 75 MHz worst intra-WNS +1.2 ns, clk_stor 27 MHz +13 ns,
MIG internals +1.46 ns) are unaffected and healthy at every candidate.
| Tier | CPU clock | Evidence | Margin | Required before shipping |
|---|---|---|---|---|
| Conservative | 25.0 MHz (clk 25) |
run_clk25: CPU WNS +9.293 of 40 ns | 23% of period - absorbs corner/aging/spread and the untimed-loop floor comfortably | SINTRAN boot + make test-instr-class validation on the board at 25 MHz |
| Balanced | 33.3 MHz (clk 33) |
run_clk33: CPU WNS +1.282 of 30 ns | ~4% of period | same board validation + a small directive/seed sweep to check spread |
| Aggressive | 40.0-45.45 MHz (clk 40/clk 45), or 50 MHz with physopt |
run_clk40 +0.319, run_clk42 +0.152, run_clk45 +0.085, run_clk50_1 +0.007 | sub-0.5 ns - no engineering margin; single-seed results | NOT recommended for unattended use or anything writing to the SD card until: multi-seed sweep shows consistent closure, the CGA IDB ring is cut in RTL (making WNS trustworthy), and a long soak passes on the board |
Functional caveats that STA cannot cover (all inherited, none created by this campaign):
- The 2 auto false paths / 6 combinational loops (CGA IDB ring) are untimed at every frequency. Cutting the ring in RTL is the prerequisite for trusting the aggressive tier.
- CPU:ui_clk ratio changes (16.667:75 -> up to 50:75). The MEM_HOLD freeze and the nds_sync toggle handshakes are contract-based (ratio-independent by design) and the payload bounds now track the CPU period, but this has never been exercised on silicon above 16.667 MHz.
BOARD_CLK_FREQmoves withclk=automatically (UART baud, RTC tick, watchdogs) - verified to be a single pair inbuild.tcl.
Still open from this campaign (checked 28-SEP-2026): the ASYNC_REG hygiene
(section 7 item 2 - nd120_nexys4ddr_top.v still has no ASYNC_REG), a
directive/seed sweep to measure spread at the deployed clock (section 7 item
3), and the CGA IDB ring cut (Verilog/docs/HANDOFF-cga-idb-ring-cut.md). The
tier choice is made: the board is deployed at 33.333 MHz with the cache ON
(../timing.md).