A0 at the UART, and a RET that fetches 0xFF
bigxfer is the serial soak test. It’s a 2233-byte program: 185 bytes of code and a 2048-byte table of random data. It goes into RAM over the serial line through the monitor’s L command, which is about 4400 hex characters, each one echoed and checked, plus a sum at the end. Then G runs it, and it streams every byte of itself back out as hex, prints a 16-bit sum and an xor, and returns to the monitor. The host compares both directions byte for byte. It’s the test the serial card has to pass before interrupts go in.
Yesterday ended with received bytes coming back with 1 bits turned to 0, about three characters in a thousand. Over a 4400-character load that is a dozen bad characters, and the loader gives up after three tries. Today that stopped, with a 22 pF cap on the UART’s A0 pin. bigxfer then went in first try and streamed all 2233 bytes back out with matching sums, and then the machine halted instead of returning to the monitor. Most of the day went to that halt. It’s a RET fetched from RAM that reads 0xFF, but only after certain instructions, and I don’t know why yet.
SIN carried the right byte
I wrote a small script that types hex digits into the monitor’s L command one at a time and stops on the first wrong echo, so the scope is left holding the bad character. With the scope on SIN (U103.10) it caught a 7 that came back as 0x00:

Marks sat at 3.0 to 3.8 V and spaces under 0.4 V. The bit edges landed within 0.2% of 9600 baud. The rails at U103 read 5.0 V and 16 mV on GND. Every tied input on the 16550 beeped to the right rail. So the wire was fine, and the byte went bad somewhere between the pin and the CPU.
One side note: SIN idles at about 2.8 V between characters. The FTDI cable is a 3.3 V part. In-frame it’s fine, so I left it.
A probe on A0 made the errors stop
I put four probes on the UART to set up a pattern trigger and ran a speed sweep, 1500 characters at each gap from 100 ms down to nothing. 10,500 characters went through with zero errors. At yesterday’s rate that should have been about thirty. The probes were curing it.
Then it was one probe at a time. Each run is 3000 characters typed as fast as the echo allows:
| What was on U103 | Result |
|---|---|
| probes on .1 .26 .28 .21 | 10,500 chars, 0 wrong |
| .21 (~RD) lifted | 3000, 0 wrong |
| nothing but SIN | failed at char 621, D 0x44 -> @ 0x40 |
| A0 (.28) only | 3000, 0 wrong |
| A0 off again | failed at char 61, D -> 0x00 |
| probe on M0 at the source end (U54.18) | failed at char 87 |
| probe ground clip only | failed at char 378 |
| M0-M2 rerouted away from everything, no probe | failed at char 9 |
| rerouted, probe on .28 | 3000, 0 wrong |
| 22 pF from U103.28 to GND | 3000 + 10,000, 0 wrong |
With the probe tip on U103.28 it’s 25,500 characters and no errors. With nothing there, eight runs out of eight failed.
This isn’t the RESET problem from 09-06. Back then a probe anywhere on the net fixed it and moving the wire fixed it. Here a probe at the other end of M0 does nothing, and rerouting the wires made it worse. Whatever it is lives at U103’s end. The 16550 has ~ADS tied low, so its address inputs go straight to the register select with no latch, and I think M0 moves at the chip before the CPU’s latch closes at the end of the read. I haven’t measured that.
I also guessed the wrong register was answering. IER reads 0x00, and a lot of bad bytes were 0x00. I wrote IER to 0x0F and read it back, and the bad byte was still 0x00. So that guess is out.
The 22 pF stays in as a stand-in for the probe until I know what it’s covering for.
bigxfer made it through and then didn’t come back
With the cap in, bigxfer loaded first try and streamed all 2233 bytes back with S 3329 X 13, the same sums the host computed. Then no prompt, three runs in a row. On the scope ~RD was flat and T0 was frozen high, so the machine was sitting on a HALT row. U61 turns out to be an inverter feeding the T-counter’s CET, not a latch, so “halted” means the microcode is parked on a row with the HALT bit set.
I blamed the cap first, since M0 is also address bit 0 for ROM and RAM. That was wrong. A four-byte program, CALL CRLF; RET, halts with the cap in and with it out. LDAI; OUT; NOP; RET returns either way.
RET fetched from RAM reads 0xFF after some instructions
This is the same family as the 09-08 “OUT before a RAM fetch” log, but wider. I wrote a pile of two-instruction RAM programs, each one something followed by RET:
| Instruction before the RAM RET | Ends at | Result |
|---|---|---|
| NOP, LDAI | T1 | returns |
| HLSP | T2 | returns |
| POPA | T4 | returns |
| JMP | T3 | returns (twice) |
| OUT (with the settle row that’s burned in) | T3 | halts |
| LDA | T3 | halts |
| MVI | T5 | halts |
| RET from ROM | T13 | halts |
The oracle can’t see any of this. It has no timing, so it returns from all of them. I added a trace option to it and a test that holds these bench results as truth, so any rule I come up with gets checked against the whole table.
I had two rules today and killed both. The first was “the instruction before ends late”, and HLSP ending at T2 and returning killed it. The second was “the T-counter clears through two or more bits”. That one correctly predicted MVI would halt before I ran it, and then JMP ended at T3 exactly like OUT and returned twice. So it isn’t just where the previous instruction ends. What its last rows do matters, and I don’t know which part yet.
A POP right after OUT reads the stack fine. An LDAM right after OUT, which uses the same C and MDR-park paths RET does, reads fine too.
What the halted machine said on the DMM
The machine is static when it halts, so a DMM sees everything. On a halted LDA 0x8200; RET with RET sitting at 0x8103:
| Register | Value |
|---|---|
| IR (U34) | 0xFF |
| PC (U1-U4) | 0x8104 |
| MAR (U54, U59) | 0x8200 |
| T | 1 |
PC is one past the RET. The fetch happened, and the byte it got was 0xFF, which is HALT. RET never ran. MAR still holds LDA’s address, so nothing else ran either. I read MAR in the wrong pin order once and spent a while on a story about RET popping the wrong address. With the pins read in order it’s 0x8200 and that story is gone. After a reset, D 8103 reads 0x34, so the RAM cell is fine.
Catching the bad fetch on the scope
That program halts every time, and the bad fetch is the last thing before HALT goes high. That makes HALT a perfect trigger.

During the RET fetch the IR latch opens, and about 60 ns later there’s a 24 ns glitch. The IR latch enable drops, the RAM buffer enable U21 spikes off, and on a second capture the RAM chip’s own ~OE spikes off at the same moment. Right after that, IR bit 0 goes to 1 and stays there until the latch closes 340 ns later. M0 stays put the whole time, so the address didn’t move. The RAM and its buffer are back on after the glitch and the cell still holds 0x34, but the IR input stays at 1.
All 256 opcodes have the same T0 row in the burned microcode, so there’s no logical path for the IR to feed back into the fetch. The EEPROM outputs can glitch while their address changes though, and the IR is transparent during the fetch, so its outputs are part of that address.

Still open
- Why IR bit 0 stays at 1 after the glitch when the RAM holds 0x34 and is driving again. The bytes reach the IR on the W bus through the U25 bridge. Next capture is MDR0, U25’s enable and U25’s DIR during the bad fetch. If MDR0 is right and the bridge lets go, it’s a control glitch on the bridge.
- Why JMP before a RAM RET is safe and OUT and LDA aren’t. All three end at T3.
- Whether the UART’s A0 problem and this one are the same thing. Both are a read into a latch that loses its 1 bits or gets an undriven bus instead.
- What the 22 pF on U103.28 is actually covering. It’s a stand-in until I know.
- Workaround to try: a NOP right before any RET that runs from RAM. Every instruction that ends at T1 has been safe so far.