Skip to content

ND-120 on Nexys 4 DDR - Timing Closure Report (clock-up campaign)

Historical / campaign closed (kept as the measured timing record). The STA numbers below stand. Since this was written, SINTRAN boots on silicon at 45.45 MHz and 50 MHz, and the board is deployed at 33.333 MHz with the cache ON (the cache added routing pressure that reopened 45.45 - see ../timing.md, sections dated 30/31-AUG-2026). Read this report for the per-domain analysis and the frequency search, not for the "next steps".

Started 26-AUG-2026, branch clock-up. All paths in this document are relative to the repo root unless stated otherwise. Evidence directories: Verilog/fpga/nexys4ddr/timing-analysis/baseline_clk16/ (preserved reports of the shipped build) and timing-analysis/run_clk<N>/ (one per implementation run, generated by the battery in build.tcl).

1. Executive summary

The shipped SINTRAN-booting build runs the CPU at 16.667 MHz with the CPU domain only ~55% utilized in time: its worst path arrives in 33 ns of a 60 ns budget. A ten-run post-route frequency search (every run a full synth->place->route with the real constraint, never synthesis arithmetic) measured the actual wall:

  • Every step from 25.0 to 45.45 MHz closes timing with the default flow. 45.45 MHz (22 ns) closes at WNS +0.085 ns; 50 MHz (20 ns) fails the default flow by -2.546 ns across 1213 endpoints.
  • Adding phys_opt_design (new opt-in -tclargs physopt) closes 50 MHz at WNS +0.007 ns - 3x the shipped clock, zero practical margin.
  • The CPU domain has exactly ONE critical-path family at every frequency: the full microcycle from the WCS microcode BRAM through IDB decode/ALU/TRAP/PALs to the register-file clock-enables, 29-30 logic levels, routing-dominated (76% at 60 ns, 64% at 22 ns).
  • One stale constraint was found and fixed (crossing bounds into the CPU clock hard-coded at 80 ns; now track the selected period). No timing exception was added anywhere.
  • STA pass is not functional proof. (When this report was written nothing above 16.667 MHz had booted; since then 45.45 MHz and 50 MHz have booted on silicon, and the deployed build is 33.333 MHz with the cache ON.) The two auto-inserted loop-breaking false paths (CGA IDB ring remnants) make every CPU WNS a floor; and WNS +0.007 (50 MHz) or +0.085 (45 MHz) leaves no allowance for the untimed paths, multi-run spread, or anything else. The defensible recommendation is below (section 9).

Recommended: 25 MHz as the new proven-margin default candidate (WNS +9.3 ns, 23% margin), 33.3 MHz as the balanced target (WNS +1.28 ns, 2x the shipped clock) - each contingent on booting SINTRAN and running the instruction-verify suite on the board at that clock, which requires flashing (Ronny's call).

2. Exact board / FPGA / tool identification

Everything below is read from the project files and the routed report, not inferred:

Item Value Evidence
Board Digilent Nexys 4 DDR (Rev C) Verilog/fpga/nexys4ddr/nd120_nexys4ddr.xdc header ("Pin source of truth: Nexys-4-DDR-Master.xdc, Digilent official, Rev. C board"); DDR2 MIG core ddr2-test/ip/ddr/ddr.xci (the non-DDR Nexys 4 has cellular RAM, no DDR2)
FPGA part xc7a100tcsg324-1 (Artix-7, speed grade -1) build.tcl line 19 (set part); routed report header "Device: 7a100t-csg324, Speed File: -1 PRODUCTION 1.23 2018-06-13"
Vivado v2026.1 (win64, Build 6511674), Windows host timing-analysis/baseline_clk16/timing.rpt header
Flow Non-project in-memory batch (create_project -in_memory), synth -> opt -> place -> route in one script build.tcl
Top module nd120_nexys4ddr_top build.tcl synth_design -top; Verilog/fpga/nexys4ddr/nd120_nexys4ddr_top.v
Oscillator 100 MHz on pin E3 (clk100) nd120_nexys4ddr.xdc: create_clock -name sys_clk -period 10.000
Baseline report state Routed ("Design State: Routed"), multi-corner (Slow + Fast), pessimism removal on baseline_clk16/timing.rpt lines 1-35
Report vs sources The 25-AUG 21:05 report matches the committed sources: the last RTL/constraint commit precedes the 21:02 build, and later commits (ebe018d, e20678d, 46b5e32) are documentation/submodule only git log
Build config of the baseline clk 16 ilaslim : 16.667 MHz CPU clock, DDR2 main RAM (MAIN_RAM_DDR2), CPU cache off (ND120_NO_CACHE), WCS preloaded (SKIP_WCS_LOAD), FF mode, slim ILA + dbg_hub in the netlist build.tcl defaults + the shipped nd120_nexys4ddr.ltx existing

3. Clock architecture

One MMCME2_BASE in nd120_nexys4ddr_top.v (line ~108), VCO = 100 MHz x 10 = 1000 MHz:

Clock Source Period / freq Moves with clk=? Drives
sys_clk pin E3 10 ns / 100 MHz no MMCM input, 7-seg + LED stretchers, heartbeat
clk_cpu_pre -> BUFG clk_cpu MMCM CLKOUT0, CLKOUT0_DIVIDE_F = ND120_N4DDR_MMCM_DIV 60 ns / 16.667 MHz at clk=16 YES - the only one entire CPU + board + ND-BUS device chain + BRAM-cache side of main RAM
clk_stor_pre -> clk_stor MMCM CLKOUT1, /37 37 ns / 27.027 MHz no SD/FAT storage stack
clk200_pre -> clk200 MMCM CLKOUT2, /5 5 ns / 200 MHz no DDR2 MIG reference
clk_pll_i (ui_clk) inside MIG, from clk200 13.333 ns / 75 MHz no DDR2 user interface, nd_ddr2_port/arb/cache backend
MIG internal (freq_refclk, mem_refclk, oserdes*, iserdes*, sync_pulse) MIG PLL 300/600 MHz etc. no DDR2 PHY
dbg_hub TCK JTAG 33 ns no ILA/dbg_hub only

The CPU period in ns equals the divider value (VCO 1000 MHz), so the clk_table in build.tcl maps clk=16 -> 60 ns, 20 -> 50, 25 -> 40, 27 -> 37, 33 -> 30. BOARD_CLK_FREQ moves with the divider so UART baud, RTC tick and watchdog counts stay correct - the two must move together (build.tcl comment, verified in the define list).

4. Baseline post-route timing (clk=16, shipped SINTRAN-booting build)

From baseline_clk16/timing.rpt (routed, 25-AUG-2026 21:05):

Metric Value
WNS (global) +1.460 ns - held by a MIG-internal path sync_pulse -> mem_refclk (report line 224), NOT by the CPU domain
TNS 0.000, 0 failing endpoints of 53987
WHS +0.024 ns (floppy DMA dma_addr_reg -> s_addr_reg, intra-CPU; hold is period-independent)
WPWS +0.264 ns (clk200 MIG domain)
clk_cpu_pre intra-domain WNS +27.014 ns of a 60 ns requirement
clk_stor_pre intra-domain WNS +13.489 ns of 37 ns
clk_pll_i intra-domain WNS +5.993 ns of 13.333 ns
Recovery (async CLR deassert, "async_default" cpu->cpu) +48.541 ns - reset-recovery checks, healthy

Worst CPU-domain path (slack +27.014, i.e. 32.986 ns arrival): CPU/CS/WCS/CHIP_21C RAMB18 (microinstruction word) -> CSIDBS decode -> INTR/CNTLR/IRGEL -> ALU/OUTMUX + STS + RALU carry -> FIDBO -> TRAP/BRKDET -> PAL_44307_UCYCLK MAP_n -> MIC/MIC_IPOS MA -> CS/ACAL LUA -> PAL_44403_UCYIN0 -> TERM_D -> ALUCLK_EN (fo=238) -> WRF/RBLOCK/Z_REG_0 clock-enable. 30 logic levels, 7.7 ns logic / 24.7 ns routing (76%) - a routing-dominated path placed under a loose 60 ns budget. This is the machine's full microcycle: microcode word out of the WCS to the write-enable of the register file.

Estimated CPU-domain boundary from this single number: 60 - 27.014 = 32.99 ns -> ~30.3 MHz. This is a timing-model estimate at THIS placement; with 76% of the path being routing under no pressure, a tighter constraint may do better - or uncover harder paths. Only reruns at candidate constraints decide (section 8).

Known caveat that makes every CPU WNS a floor, not a guarantee: synthesis auto-inserts two loop-breaking false paths through the historical CGA IDB ring (ALU_OUTMUX/D_15_0[8] and ALU_i_426/O, Synth 8-326; check_timing reports 6 combinational loops). Paths through those nodes are untimed. Inherited, documented in build.tcl ("PARKED DEBT"); the RTL ring cut is the real fix and is out of scope for this campaign.

5. Constraint audit

# Finding Verdict / action
1 create_clock on clk100 (sys_clk, 10 ns) present; all MMCM/MIG generated clocks auto-derived (check_timing generated_clocks: 0) OK
2 set_max_delay -datapath_only bounds INTO clk_cpu were hard-coded 80.000 ns (one destination period of the 12.5 MHz era) while the actual period is 60 ns and shrinking. Contract in the comment: "no payload bit may take longer than one destination period" FIXED 26-AUG in build.tcl: bound = $mmcm_div (the CPU period in ns), tracks clk=. Safe direction - strictly tighter
3 nd120_timing.xdc set_clock_groups -asynchronous (sys_clk vs clk_cpu_pre) OUTRANKS the later set_max_delay cpu->sys 10 ns in build.tcl, so cpu->sys crossings are untimed (no cpu_pre->sys row in the Inter Clock Table) Accepted: those crossings are single-bit flags through 2-FF synchronizers (nd120_nexys4ddr_top.v sync_hold/sync_grb), where max-delay is a nicety, not a correctness need. Documented, not changed
4 check_timing: 5 no_input_delay (UART RX, SD card inputs, buttons), 40 no_output_delay (LEDs, 7-seg, UART TX, SD) Accepted: asynchronous human/serial interfaces; SD bus is source-synchronous to the FPGA-driven sd_clk in the fixed 27 MHz storage domain, unchanged by this campaign
5 6 combinational loops + 2 auto false paths (CGA IDB ring remnants) Inherited parked debt, see section 4 caveat
6 Reset recovery/removal: timed (async_default recovery checks), worst arrival 10.6 ns OK at every candidate period
7 No multicycle paths declared anywhere Correct as-is: no path has been PROVEN multicycle by RTL analysis. None added (rule: no exception without functional proof)
8 MIG constraints come with the core (read_ip), pins included OK, untouched

6. Critical-path families

Per-domain extraction (report_cpu_paths.tcl on the run checkpoints, setup_paths_cpu_group.rpt in each run directory):

The CPU domain has exactly ONE critical-path family. At clk=16 AND at clk=33, all 100 worst CPU-group paths are the same shape:

  • Startpoint: WCS microcode BRAM read clock (CORE/CPU_BOARD/CPU/CS/WCS/CHIP_21C..22D/idt_memory_array_reg/CLKARDCLK, RAMB18E1)
  • Endpoint: register-file clock-enables (CORE/CPU_BOARD/CPU/PROC/CGA/DELILAH/WRF/RBLOCK/<R0-R7,Z>_REG_*/regFF_reg[*]/CE, FDRE)
  • Route (from the clk=16 worst path, baseline_clk16/timing.rpt): microcode word -> CSIDBS_4_0 IDB-source decode -> INTR/CNTLR/IRGEL -> ALU/ALU_OUTMUX + ALU_STS -> ALU_RALU CARRY4 chain -> FIDBO -> TRAP/TVGEN + BRKDET -> PAL_44307_UCYCLK (BRK_n/MAP_n) -> MIC/MIC_IPOS (MA_12_0) -> CS/ACAL (LUA_12_0) -> PAL_44403_UCYIN0 -> U_TERM_D -> ALUCLK_EN (fanout 238) -> WRF CE. Functionally: this is the machine's full microcycle - the microinstruction coming out of the WCS deciding, combinationally, whether the register file clocks at the end of the same cycle.
  • Composition: 29-30 logic levels (LUT2..LUT6, 3x CARRY4), logic delay is only 7.0-7.8 ns; ROUTING dominates at every constraint tried and keeps compressing under pressure: 25.3 ns route (76%) at the 60 ns budget, 21.1 ns (75%) at 30 ns, 13.877 ns (64%) at 22 ns (measured: run_clk45/timing_summary_post_route.rpt, worst CPU path 21.681 ns = 7.804 logic + 13.877 route).
  • Root cause class: excessive combinational depth (a faithful transcription of the original board's asynchronous microcycle logic - PALs and TTL between two clocked elements), amplified by high-fanout control distribution (ALUCLK_EN fo=238, CSIDBS fo=35, MA_12_0 fo=39).

No second family surfaced anywhere in the search - when the microcycle path is met, the whole CPU domain is met. The fixed-frequency domains (clk_pll_i 75 MHz DDR2 user logic, clk_stor_pre 27 MHz SD stack, MIG PHY) hold the global WNS once the CPU domain is fast, and none of them moves with clk=.

Supporting observations from the clk=16 battery (run_clk16/):

  • qor_assessment.rpt: design "easily meets timing" at 60 ns; no ML suggestions generated.
  • high_fanout_nets.rpt: the largest fanout nets are reset distribution in the storage/CPU domains (2628, 1416, 778, 602) - all with >24 ns slack at their periods; not limiting.
  • cdc.rpt: CDC-8 x2 (reset synchronizers missing ASYNC_REG), CDC-5 x1 (multi-bit synchronizer missing ASYNC_REG) - hygiene items, listed in section 7.
  • methodology.rpt: TIMING-24 x1 confirms one datapath-only bound is overridden (the cpu->sys bound under the clock group, audit item 3); SYNTH-6 x109 (BRAM output timing sub-optimal - the WCS BRAMs read combinationally into the microcycle, no output register: that IS the critical-path family); LUTAR-1 x6 (LUT drives async reset).

7. Improvement options

The frequency search (section 8) showed the as-is RTL already closes 45.45 MHz - 2.7x the shipped clock - so no RTL change is NEEDED to raise the clock substantially. Options if more speed or more margin is wanted, ranked:

Low risk (no functional change):

  1. phys_opt_design after place and after route (-tclargs physopt, added 26-AUG as an opt-in flag). Expected: recovers a few hundred ps on routing-dominated paths; tested at clk=50 (run_clk50_1, section 8). Effort: none (done). Preserves cycle accuracy: yes.
  2. ASYNC_REG hygiene: add (* ASYNC_REG = "true" *) to the three flagged synchronizer chains (CDC-5/CDC-8 in run_clk16/cdc.rpt, incl. the debug-panel sync_hold/sync_grb in nd120_nexys4ddr_top.v). Benefit: keeps synchronizer pairs placed together (MTBF), removes the warnings; no Fmax change. Effort: minutes. Risk: none.
  3. Directive/seed sweep near the wall (place_design -directive, route_design -directive, e.g. ExtraTimingOpt / AggressiveExplore). Vivado is deterministic per configuration, so the single-seed +0.085 at 22 ns says nothing about spread; a sweep establishes "consistently achievable" vs "lucky". Effort: ~40 min per run, no code change.

Medium risk (RTL, changes netlist but not architecture):

  1. Register the WCS BRAM output (RAMB18E1 DOA register, DOA_REG=1). Removes the BRAM clock-to-out (2.45 ns) from the family head and lets the BRAM place freely - BUT the microinstruction then arrives one cycle late, which changes the microengine's fetch/execute overlap. On this design that is an ARCHITECTURAL change in disguise: NOT recommended without a full golden-trace revalidation (make compare, instruction-verify campaign).

Architectural (documented only - explicitly out of scope for a cycle-faithful reconstruction):

  1. Pipeline the microcycle (split at FIDBO or at the MIC address formation). Would roughly double Fmax and roughly halve instructions-per-clock; destroys cycle accuracy against the original machine and every validated timing assumption (RTC, UART pacing, device handshakes). Not recommended for this project's goal.
  2. Break the remaining CGA IDB combinational ring in RTL (the "PARKED DEBT" in build.tcl): removes the 2 auto false paths and the 6 check_timing loops, making the WNS a real guarantee instead of a floor. This is a correctness-of-analysis improvement, not a speed improvement, and the right long-term move (tracked in the 20-AUG analysis, ../BUILD-WARNINGS-ANALYSIS.md).

Constraint work (done during this campaign):

  1. FIXED: set_max_delay -datapath_only bounds into clk_cpu now track the CPU period ($mmcm_div) instead of a stale 80 ns (section 5, item 2).
  2. NOT done, documented: no multicycle exceptions were added anywhere. The 52 ns WCS->ACAL "multicycle" claim from the Tang notes was NOT assumed true here - no RTL proof exists that the path is multicycle, and the search closed 45 MHz without it.

8. Experiments

All runs: build.tcl flow, -tclargs clk <N> ilaslim -noburn, full post-route battery + checkpoint in timing-analysis/run_clk<N>/. Same config as the shipped SINTRAN-booting build except the CPU divider.

Run Change vs previous clk_cpu WNS global WNS TNS WHS Result
baseline (25-AUG) shipped build, 80 ns stale ->cpu bounds +27.014 (60 ns req) +1.460 0 +0.024 reference; SINTRAN boots 7/7
run_clk16 ->cpu crossing bounds now track the period (60 ns) +26.455 (60 ns req) +1.460 0 +0.012 PASS - tightened bounds cost nothing; CPU worst arrival 33.0 ns
run_clk25 CPU divider 40 -> 25.0 MHz +9.293 (40 ns req) +1.224 0 +0.032 PASS - CPU worst arrival 30.7 ns; global WNS now in fixed 75 MHz clk_pll_i domain
run_clk33 CPU divider 30 -> 33.333 MHz +1.282 (30 ns req) +1.282 0 +0.019 PASS - CPU worst arrival 28.7 ns; CPU domain is now the design-wide binding domain
run_clk35 new clk_table entry: divider 28.0 -> 35.714 MHz +1.308 (28 ns req) +1.308 0 +0.023 PASS - CPU worst arrival 26.7 ns
run_clk38 new clk_table entry: divider 26.0 -> 38.462 MHz +0.316 (26 ns req) +0.316 0 +0.013 PASS by a hair - CPU worst arrival 25.7 ns; margin collapsed, wall is here
run_clk40 new clk_table entry: divider 25.0 -> 40.0 MHz +0.319 (25 ns req) +0.319 0 +0.015 PASS - CPU worst arrival 24.7 ns; router still tracking the constraint
run_clk42 new clk_table entry: divider 24.0 -> 41.667 MHz +0.152 (24 ns req) +0.152 0 +0.025 PASS - CPU worst arrival 23.85 ns
run_clk45 new clk_table entry: divider 22.0 -> 45.455 MHz +0.085 (22 ns req) +0.085 0 +0.023 PASS - CPU worst arrival 21.9 ns
run_clk50 existing entry: divider 20.0 -> 50.0 MHz -2.546 (20 ns req) -2.546 -1600.263 / 1213 endpoints +0.020 FAIL - gate refused the bitstream; the wall is between 20 and 22 ns
run_clk50_1 + physopt (post-place & post-route phys_opt_design -directive AggressiveExplore, new opt-in flag) +0.007 (20 ns req) +0.007 0 +0.020 PASS - phys_opt recovered the full 2.55 ns; 50 MHz is STA-legal with zero practical margin. SILICON 26-AUG: SINTRAN III BOOTS at 50 MHz (this bitstream, one boot, banner reached, console driven by Ronny)
run_clk50_2 console baud 9600 -> 115200 (UART divider constant 5208 -> 434), same clk 50 + physopt -0.210 (20 ns req) -0.210 -4.586 / 54 endpoints +0.052 FAIL - the one-constant change re-rolled placement and the +0.007 closure did not survive. 50 MHz closure is FRAGILE: any netlist edit is a new seed
run_clk45_1 clk 45 + physopt + 115200 baud +0.020 (22 ns req) +0.020 0 +0.031 PASS - programmed 26-AUG 10:17, SINTRAN BOOTS on silicon with the 115200 console (Ronny-verified). Deployed configuration

All runs one seed each (Vivado is deterministic per configuration - spread across directives/seeds is NOT known; see section 9). WHS stayed in +0.012..+0.043 across all runs (hold is period-independent and healthy). WPWS +0.264 (MIG clk200 domain) at every run.

9. Frequency roadmap and remaining work

Estimated boundaries per domain (post-route, per-clock WNS, NOT one global number): the CPU domain is the only one that moves with clk=; the fixed domains (clk_pll_i 75 MHz worst intra-WNS +1.2 ns, clk_stor 27 MHz +13 ns, MIG internals +1.46 ns) are unaffected and healthy at every candidate.

Tier CPU clock Evidence Margin Required before shipping
Conservative 25.0 MHz (clk 25) run_clk25: CPU WNS +9.293 of 40 ns 23% of period - absorbs corner/aging/spread and the untimed-loop floor comfortably SINTRAN boot + make test-instr-class validation on the board at 25 MHz
Balanced 33.3 MHz (clk 33) run_clk33: CPU WNS +1.282 of 30 ns ~4% of period same board validation + a small directive/seed sweep to check spread
Aggressive 40.0-45.45 MHz (clk 40/clk 45), or 50 MHz with physopt run_clk40 +0.319, run_clk42 +0.152, run_clk45 +0.085, run_clk50_1 +0.007 sub-0.5 ns - no engineering margin; single-seed results NOT recommended for unattended use or anything writing to the SD card until: multi-seed sweep shows consistent closure, the CGA IDB ring is cut in RTL (making WNS trustworthy), and a long soak passes on the board

Functional caveats that STA cannot cover (all inherited, none created by this campaign):

  1. The 2 auto false paths / 6 combinational loops (CGA IDB ring) are untimed at every frequency. Cutting the ring in RTL is the prerequisite for trusting the aggressive tier.
  2. CPU:ui_clk ratio changes (16.667:75 -> up to 50:75). The MEM_HOLD freeze and the nds_sync toggle handshakes are contract-based (ratio-independent by design) and the payload bounds now track the CPU period, but this has never been exercised on silicon above 16.667 MHz.
  3. BOARD_CLK_FREQ moves with clk= automatically (UART baud, RTC tick, watchdogs) - verified to be a single pair in build.tcl.

Still open from this campaign (checked 28-SEP-2026): the ASYNC_REG hygiene (section 7 item 2 - nd120_nexys4ddr_top.v still has no ASYNC_REG), a directive/seed sweep to measure spread at the deployed clock (section 7 item 3), and the CGA IDB ring cut (Verilog/docs/HANDOFF-cga-idb-ring-cut.md). The tier choice is made: the board is deployed at 33.333 MHz with the cache ON (../timing.md).