Add-with-carry is one more '157

/ DINO / updated 2026-10-06 / from dino-homebrew @ 6d1a945 dino alu microcode 74ls157 mandelbrot

DINO has add-with-carry. Four new instructions, ADC, SBB, ACI and SBI, ran clean on the bench at 1.024 MHz, and the cost in hardware was one 74LS157 and one wire moved to it. I’d honestly assumed ADC was already in there. It wasn’t. The microcode bit for it has been sitting unused on the third EEPROM since August, waiting for this chip.

ImageResetsOB every time
isa (all 147 old subtests)100xB4
isasoak100x00
carry (new)100xAD

How the carry gets into the ALU

The carry-in pin (CN) of the low 74F382 used to come straight from a NAND of two ALU select bits, so ADD always got 0 and SUB always got 1. U76 is a ‘157 that picks between that old signal and the carry flag:

U76 pinNetFrom
1 SCW16U23.11, the third microcode EEPROM
2 I0aFLAG_CU49.2, the flag register
3 I1aALU_CINU50.13, the old NAND
4 ZaALU_CNU38.15, the ‘382’s CN
15 EGND

Every old microcode row has CW16 high, so the mux passes the old carry and the machine is unchanged. The four new opcodes pull CW16 low on their ALU row and get the flag instead. One flag works for both directions because on the ‘382 CN is a true carry when adding and a not-borrow when subtracting, so ADC and SBB both want CN = FLAG_C with no inverter.

The only existing wire that moved is the one on U38.15. Everything else was new.

isa read 0x01 because pin 15 was loose

First power-up with the chip in and the old microcode, isa read 0x01. That’s the first subtest, a load-immediate that gets checked with a compare. A compare is a subtract, and a subtract needs CN = 1.

U76.15 is the ‘157’s active-low enable. It wasn’t seated, a floating TTL input reads high, and a ‘157 with E high forces every output low. So CN sat at 0, every subtract came out one low, and the very first compare in the program failed. ADD still worked, because ADD wants CN = 0 anyway.

If a 74LS157 output is stuck low and the select pin looks fine, check that pin 15 (E) is actually grounded.

Pushed it in, isa read 0xB4, and isasoak gave 0x00 ten times. That proved the mux was invisible before any new microcode went in.

The new opcodes

OpcodeInstructionDoes
0x96ADCA = A + B + C
0x97SBBA = A - B - borrow
0x98ACI nA = A + n + C
0x99SBI nA = A - n - borrow

ACI and SBI reuse ADI and SUI’s rows exactly and only clear CW16 on the last one. The operand fetch before it doesn’t touch the flags, so the carry they read is whatever the previous instruction left. Every other row in all three EEPROMs is byte-identical to the image that was in the sockets, and a host test pins that with a CRC.

The microcode generator now refuses a carry-from-flag row on a logic op or without the ALU as the source. Both would encode cleanly, burn, and quietly behave like the same row without the bit. The Python oracle refuses an ADC that follows an AND, since the ‘382 datasheet doesn’t define the carry after a logic function.

PROG_carry runs 12 subtests 256 times and stops on the first failure. Each op has at least one case where the carry flips across it, carry in 1 and carry out 0 or the other way around. The flag register clocks the new carry on the same CLK edge the ALU’s output latch closes on, and an ALU that grabbed the new carry would read one off on exactly those cases. It didn’t, 10 runs out of 10.

bigxfer’s checksum

bigxfer streams 2K of itself back over serial with a 16-bit sum. The sum used to be ADD, then JNC over an INR of the high byte. It’s ADD, then ACI 0 on the high byte now, with no branch. The carry has to survive a load and a store between the two, and it does. It’s the first RAM program using the new instructions, and it ran fine.

Mandelbrot in 8.8

The old mandel.asm keeps every number in one byte, 1/32 units, because there was no way to carry into a second byte. mandel88.asm is 16-bit signed 8.8 fixed point throughout, 1/256 resolution, with every add and subtract done as ADD/ADC or SUB/SBB pairs.

There’s still no multiply. Mandelbrot only needs squares, because 2xy = (x+y)² - x² - y². A square comes out of one 256-byte table built at startup, T[l] = floor(l²/256). For a magnitude n = 256h + l:

n²/256 = 256·h² + 2·h·l + l²/256

The first two terms are whole numbers, so floor(n²/256) is exactly 256·h² + 2·h·l + T[l]. The 2hl part is just adding l to a 16-bit sum 2h times, and h is at most 3 here. The table gets built with a 16-bit running sum, n² = (n-1)² + n + (n-1), which is ADD and ACI again.

The window lives in four 16-bit RAM cells at 0x811C, so the picture zooms with W commands from the monitor prompt and no reload. The default window draws the same grid as the old program, and it matched what the oracle computed on every row I checked by eye. Against floating point it disagrees on 14 of 2640 cells. At 8x zoom the grid step equals the arithmetic’s own resolution, and that goes to 45 of 2640.

The default picture takes 22 s on the bench, twice the same. The oracle counts 22.4 million T-states of compute for it, which is 21.9 s at 1.024 MHz, so the serial wire adds nothing you can see. Each character takes around 8 ms to work out and 1 ms to send, so PUTC never has to wait. I’d guessed 27 or 28 s by scaling the old program’s 14 s, and that guess was wrong. The old one’s own count is 8.6 million T-states, 8.4 s, and I can’t account for the other 5 or so seconds yet.

mandel88.asm, default window, on the serial terminal at 9600 baud. Same grid as the one-byte version, now in 16-bit 8.8 fixed point. 22 s at 1.024 MHz.
mandel88.asm, default window, on the serial terminal at 9600 baud. Same grid as the one-byte version, now in 16-bit 8.8 fixed point. 22 s at 1.024 MHz.

Six W commands from the prompt moved the window to x 0.28 to 0.59, y 0.56 to 0.31, the edge of the main cardioid above elephant valley, at 1/256 across and 1/128 down. That’s 8x the default, and the old program has 10 columns across that whole area. The rows I compared against the oracle match:

@@@@@@@@@@@@@@@@@@#@*++====xx#@+-;;:::::::,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,.......
@@@@@@@@@@@@@@@@@@#@====-==--===;;::::::::,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,.......
@@@@@@@@@@@@@@@@@@*+==-----;;;;;;::::::::::,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,.....
8x zoom, x 0.28 to 0.59 and y 0.56 to 0.31, at 1/256 across and 1/128 down. The one-byte version has 10 columns across this whole area. Also 22 s.
8x zoom, x 0.28 to 0.59 and y 0.56 to 0.31, at 1/256 across and 1/128 down. The one-byte version has 10 columns across this whole area. Also 22 s.

The zoom took 22 s too. Zooming doesn’t change what one iteration costs, so the time follows the total iteration count, and the two windows land within 1% of each other: 30,439 iterations for the default and 30,264 for the zoom, from the model. The zoom has fewer cells inside the set, 676 against 884, but its outside cells sit close to the edge and take longer to escape. A window that was all inside would be 63,360 iterations, somewhere around 45 s.

Ctrl-C under imon broke out of a picture partway through, same as it does for anything else.

Open threads