The plan, reordered: known bugs before new functionality¶
Standing rule, set 2026-08-18: known bugs are ALWAYS fixed before new functionality. This supersedes the ordering implied by the chat design docs, which had features queued ahead of an open defect.
Every phase ends with something proved on a real machine, not argued. Numbers in brackets are task ids.
Phase 0 - restore the lab [53]¶
Nothing live can be done until this is green. The third failed close experiment killed *FA-SERVER
and *FA-FSA on D100; the code was reverted, the machine was never re-verified.
X-C: LIST-NAMES- expect one*FA-SERVERwith free seats and one*FA-FSA.- If missing:
ABORT FSART,RT FSARTfirst, thenFS-ADMINISTRATOR/SELECT-FSA/START-SERVER 1,,,,. The ABORT/RT pair must come first when*FA-FSAitself is down, or the administrator answers "Remote FSA is not running" and can restart nothing. - Clear terminal 9, unresponsive after an earlier
STOP-TERMINAL.
Proof: one small push exits 0 and the seat count moves by exactly one.
Phase 1 - the FA connection-seat leak [54] DONE¶
SOLVED 2026-08-18, commits a6e8f047 and ec7f0772. Three things had to be right together,
which is why every partial attempt was useless or destructive:
| wrong | right | |
|---|---|---|
| type | 07C0 Close - the SERVER's message |
0782 Release - the client's |
| operands | LetterEchoWord first |
sender's conversation first, then the peer's |
| flags | 0x96 / 0x00 |
0x82 / 0x84 |
FaReadDriver had the identical bug. And a refused transfer leaked too - NextAction tested the
failure first so the ladder ended before SendClose; a refusal now owes a Release sent outside the
ladder.
Measured: 3 pushes + 2 pulls with *FA-SERVER steady at 29, then 3 refused creates + 2 refused
pulls + 1 good pull - zero seats spent. Tests in FaClientReleaseFrameTests.cs;
FaWriteDriverTests widened, since naming FaMessageType.Close exactly had made it stop seeing the
teardown.
Not established, and blocking nothing: the routine in FA-SERVER-TAD (segment 73B) that calls
XMPINFC. The FA server is a COSMOS product and is not in the NPL sources - XMPINFC,
FA-SERVER and QFORM appear nowhere in either tree; XSNSP only ever in symbol tables
(=000121 in J, K03 and L07). The investigation brief is retired:
E:\Dev\Ronny\NDInsight\SINTRAN\XMSG\DOC\REQUEST-FA-SEAT-LEAK-INVESTIGATION.md.
Phase 2 - our own seat and name leaks [60, 61]¶
Same family, in our code, far cheaper. Both start with a measurement, not a fix.
[60] A vanished client's seat is never freed. CHATSV returns a seat only on a clean /leave.
A crashed client holds its seat for good.
Both verify-first questions are answered, and they moved the fix off XMTRE entirely.
- Does
sendTosetXFSEC? No -xmpsend(0, ...), flags zero. So an undeliverable message is "discarded and released by XMSG" and there is nothing to react to. Read from the code. - Does a dead port really produce
XMTRE? Yes, but only for a SECURE message: XFSEC returns it "if the receiving port is closed (e.g., if the receiving task terminates)". So the port case is covered, not just the unreachable system.
But XMTRE turned out to be the wrong signal anyway, and that is the finding. A returned message
carries what WE wrote - a Said naming the speaker - so it does not say who it failed to reach. The
recipient is exactly the field a bounce does not have.
A synchronous status looked like it would have it, and does not. XMPSEND is given a magic
number and a magic carries a generation, so it seemed XMSG would refuse one naming a closed port on
the spot with XMXEIMA 16915. That was built, pushed, compiled and run.
MEASURED, and refuted: two clients in the room, one killed with ESC (USER BREAK AT 17202B), the
other spoke. The broadcast to the dead port returned XMOK. Nothing printed, the seat stayed
spent, LIST-NAMES held at 14. XMPSEND does not validate the destination.
So XFSEC after all - and the doubt about it was misplaced. With the flag set the bounce comes
back as XMTRE, and XMPFMST on it hands back the port that could not be reached, not us. That
was the open question and the machine answered it:
SV: returned msg, st= 16915 <- XMXEIMA, as the REASON ON A RETURN
SV: it names slot 2 <- the bounce identifies its destination
SV: reaped dead slot 2
16915 is the same code the refuted attempt watched for - right number, wrong place. It is never the
status of the send; it is the reason on the return.
Proved: CHAT-LOBBY 16 free, 14 with both joined, back to 15 once the dead member was
reaped, with the surviving member's seat still correctly held.
Marking and acting are kept apart: sendTo runs inside broadcast's loop and there is one shared
outBuf, so building a departure notice at that depth would overwrite the message the room is still
being sent. The mark is reaped at the top of the receive loop.
[61] REFUTED - there is no defect here. The premise was that CHATSV never closes its port, so the name lingers. Both halves are wrong:
- SINTRAN disconnects automatically "on return to the SINTRAN command processor" and "on log out or RT program termination", and a disconnect closes every port the task opened (ND-60.164.3, XMPFDCT);
- closing a port clears its name: "If the port had a name, the name is cleared (i.e., the name is removed from XROUT's name table)" (XMPFCLS).
MEASURED 2026-08-18. LIST-NAMES with the server running showed 100 9 15 CHAT-LOBBY. The
program was broken out of (USER BREAK AT 10164B) and returned to @. LIST-NAMES immediately
after: CHAT-LOBBY is gone.
So the lingering seen during [55] was not the program failing to clean up - it was the program never
TERMINATING, because the terminal had been taken out from under it with STOP-TERMINAL and the task
stayed alive holding the port. That is an operating procedure, not a bug, and the procedure is:
break the server out of its loop and let it reach @. Nothing to fix; no shutdown path needed.
Phase 3 - the small unknowns [56, 57]¶
Neither blocks anything; both are cheap and stop future confusion.
- [56] The daemon's remote-existence note is in-memory, so a restart costs one refused create per file it has never carried. Self-healing and bounded - a deliberate choice to avoid a ledger-file format change. Revisit if restarts become frequent.
- [57]
A2 4104, the follow-on refusal after a refused open - not 46, matches no ND error number, identical for two different missing names, so a fixed code with no known meaning.XRMFL, which refuses while the message table shows 4 of 256 in use, and clears on a plain retry.
Phase 4 - prove what is already committed [55]¶
The 16-seat CHATSV is committed source that no ND has ever run. Not a defect, but not proven either.
Push, build, check the listing for *** ERROR - the second pass's "0 DIAGNOSTICS" sits happily
under a compile that had three - then confirm LIST-NAMES shows 16 free seats.
Phase 5 onward - new functionality¶
Only once the phases above are green.
- [59]
CHAT:CNFGand alias memory. Chosen as the first feature. No server change, no trunk, provable on one machine. Verified feasible: the runtime provides MON 50/54/1/2 directly. Half done already, and not as a feature - as a prerequisite. The hard-codedRONNYturned out to block the [60] measurement outright: two of these clients could never be in one room, because the second one's join is refused with "that nickname is taken". So the client now ASKS at start-up (askNickname), andmyNameis a sixteen-byte buffer instead of a five-byte literal - writing a longer name into the literal would have run straight past the end of it. What remains is only the REMEMBERING:CHAT:CNFGin the user's own file area. - Local direct messages -
@and/direct, the server-scope member table, delivery reports naming the machine,/whoand/machines. - [58] Federation - phases A to E: RT server and front end, one trunk, federate state, federate messages with relay, then resilience.
- Remote direct messages and friend presence, after federation phase C.
What changed in this ordering, and why¶
The chat design docs put CHAT:CNFG first because it is small, self-contained and useful. That is
still true, and it is still the first feature. But an open defect that silently breaks unattended
file transfer outranks a convenience, and the seat leak has already cost this project days under the
name "random stalls".
One thing worth saying plainly: none of the chat work depends on the FA leak being fixed. The trunk design deliberately avoids connection seats entirely by using a normal named port. So this ordering is a discipline, not a dependency - which is exactly why it needs to be written down, or it will quietly be ignored the next time something more interesting turns up.
Corrections folded in¶
MAC is installed on D100. A comment in CHAT.PLNC claimed otherwise and used it to close off the
manual's route to any monitor call the runtime does not wrap. Checked: @MAC answers - MAC -. It is
a reentrant subsystem, so LIST-FILES finding no :PROG file proved nothing. TNOWAIT (307B) is
therefore reachable with an interface routine assembled in MAC and loaded before the PLANC library,
and the client's busy loop is a choice rather than a necessity.
Three tasks were marked complete that were not done - the lab restore, the seat leak, and the 16-seat build. Corrected. A task list that lies is worse than no task list.