Saturday, September 26, 2026

Why just switcharoo'ing RAM on a RPI isn't guaranteed to work, or "wait, you have to do WHAT to get DDR to work?"

There's been some fun stuff lately in the maker community around Raspberry Pi and RAM upgrades. Specifically, a recentish (~ 2025) boot firmware upgrade locks the RAM size to what it shipped with. There's been some write-ups about it - https://www.jeffgeerling.com/blog/2026/raspberry-pi-ram-lockdown/ and https://hackaday.com/2026/09/22/raspberry-pi-ram-restrictions-no-big-deal-frankly/ come to mind - but they don't REALLY go into all the gory details behind the scenes.

So let's do a quick dive into RAM and how it's changed from the past until now.

First off - a lot of tinkering in retro computing exposes people to static and dynamic RAM chips, first dating back to the late 70s. After a few attempts at stuff, manufacturers soon ended up kind of iterating towards a few standard pinouts and specification timings - likely so new upstart RAM makers could sell their chips as "drop in replacements" for existing RAM.

This led manufacturers towards defining an actual standard for this stuff - JEDEC (https://en.wikipedia.org/wiki/JEDEC_memory_standards) - which defines what and how various types of RAM need to behave. This is needed so you didn't end up with RAM chips needing to be paired with very specific controllers / CPUs. Some vendors kept the market on their toes with very not compatible stuff (think the RDRAM from Rambus - https://en.wikipedia.org/wiki/RDRAM) but in general one of the big reasons we had affordable RAM until recently (and will again, I'm sure) is because there's basically no vendor lock-in and a lot of interoperability.

However, that interoperability comes at a cost. Silicon isn't just some magical thing you can stamp out and get a million chips that all behave exactly the same. There's always going to be subtle behavioural differences across them. For CPUs this can actually show up as different maximum clock frequency versus heat dissipation - some CPUs will run at the maximum clock frequency at a lower temperature versus others - and yes when you graph them you'll likely get a normal or bi-modal distribution. For desktop CPUs with active cooling this may result in the same CPU being sold at a higher price with a higher maximum clock frequency, and then a "slightly cheaper" CPU being sold with a lower maximum clock frequency. If it's going into something like a cellphone or a piece of wearable technology where there's no active cooling, you either have to choose whether to lower EVERYONES maximum device performance, allow SOME devices to operate at a maximum higher performance, or you reject a percentage of the CPUs from the product and use them elsewhere.

(Yes, this is what they call "binning".)

For RAM, stable frequency, timing and bus characteristics are critical. Remember these parts aren't operating at 1 or 10MHz anymore - these chips operate their busses on both sides of the clock cycle at over 1GHz. At that point you're well into the "this behaves like RF" domain. As part of this, the PCB layout design, the connector design, the individual track lengths (think phase delays) + lengths inside the memory controller all have to line up exactly correctly - and then you still have weird RF style behaviour going on. Remember how I said every chip behaves differently? They're going to have slightly different timing and slightly different RF characteristics. So you can't just whack a single set of parameters down and hope they work - you need to go through calibration, validation and bulk testing before you know you've got something that works.

First up - the DDR4 power-up sequence is not just "apply power, add RAS/CAS timings and go." There's a whole calibration sequence that has to happen to calibrate the memory controller PHY (physical interface). For an example of this, take a look at the AMD Soft core (FPGA) DDR4 SDRAM controller calibration / setup documentation at https://docs.amd.com/r/en-US/pg353-versal-acap-soft-ddr4-mem-ip/Memory-Initialization-and-Calibration-Sequence . It has to be able to run this sequence to completion, every time, or you won't have reliably working RAM. 

What parameters are required to make this work?

Well, we get cheap, mostly interoperable DIMMs because the JEDEC standard defines what parameters to put into the DIMM EEPROM module that is read by the memory controller during power on. Yes, it's also documented online - https://en.wikipedia.org/wiki/Serial_presence_detect has all the fun details. Note that there's many, many more parameters here than just RAS/CAS timings! There's all kinds of timings to do with voltage calibration management, nanosecond accurate wait and hold times for various parts of the RAS/CAS signaling, data line skew, and other fun behavioural things.

Now here comes the fun part.

A DIMM manufacturer is going to make some test DIMMs and test them on a whole variety of motherboards to make sure that the parameters they program in reliably result in the above PHY calibration/setup succeeding 100% of the time on all the hardware they care about. They're then going to choose (hopefully conservatively) suitable values to write into the SPD to allow this to, yes, calibrate/setup 100% of the time. They may even slow down the RAM to meet timing if needed. But typically they're going to hope that the manufacturer binning of the SDRAM parts into specific part numbers and performance/behavioural bands will result in being able to make a DIMM that consistently operates at a specific performance tier without them having to further bin things when they're on the DIMM. Noone wants to throw away manufactured things, so the manufacturer has the option to just write slower parameters into SPD and sell the DIMM as a slower DIMM.

But an embedded board developer doesn't have some EEPROM attached to each group of RAM modules they've put on the board. They will have likely characterised one, two or three RAM module suppliers (either for different sizes, or the same size but multiple manufacturers so they can play off pricing against each other when making different batches of boards) and put those parameters into a fixed lookup table on said embedded board as a starting point. If you're fancy (and I don't know if RPI is being fancy or not) you'll actually run the calibration on your manufacturing line and write the SDRAM parameters for your chosen RAM into the EEPROM as the "active" configuration.

And here you see the dilemma.

See, the configuration parameters are either conservative ones for a handful of tested RAM chips from specific manufacturers, or they're very specifically for the RAM you've had installed on your embedded board. They've done a whole lot of characterization of the RAM across first one, then three, then a dozen or two, then hundreds, and finally the whole run as they're making them, to minimise the chances that during actual bulk manufacturing that a board fails RAM calibration/setup.

So when you just switch the SDRAM chip over to some other manufacturer part? You're basically hoping that given the initial RAM configuration parameters stored in EEPROM/flash somewhere on your board, the RAM calibration/setup will always succeed and then normal operation is stable. And you may get lucky! You may not! But you absolutely aren't guaranteed anything at this point.

And this is why the Raspberry Pi Foundation is doing the right thing here. Sorry y'all, but at this point you're actually likely not qualified to do all of the above work.

... and why not? Why can't you calibrate/test the RAM and write the parameters back?

A great question! The short answer is this.

It's typically vendor provided tooling that is not open source and is very proprietary, and they don't want you to know how to do it or what their memory controller is doing.

And THAT is a discussion for another time.

So now you know. It's not as simple as "RPI foundation is good" or "RPI foundation is bad". It's an intersection of how modern hardware works, how much effort goes into making sure it's bullet proof (or as bullet proof as they can get it) and what vendors / manufacturing lines are willing to release versus what they're requiring to keep as closed / NDAed stuff.

 

(Don't hate the player, hate the game, etc) 

Sunday, June 7, 2026

The great 2024/2025/2026 tech firings, and what happens to all that clue

So yeah, I was caught up in the Meta/Facebook layoffs in March 2026.

There are a lot of parallels to the 2001 dot com bust, and I'd like to just briefly touch on one of them now.

There's been a lot of layoffs over the last year or three in tech. I think the tech market has seen what, more than 100,000 in the last year alone? And most of those aren't due to performance. There are a lot of very smart people who've spent quite a few years learning how to build, debug and maintain all kinds of interesting stuff.

And they're now in the marketplace.

When this happened in 2000/2001, you had a whole lot of people who learned how to build internet tech and infrastructure at companies that were pushing the boundaries with things. And they were let go, due to downturns, change of focus, all kinds of reasons.

A lot of those people went off to eventually build the new stuff that likely overtook a lot of those internet and telco companies in the 90s.

I think this is going to happen with the layoffs from Meta/Facebook and other technology companies. Meta ran (runs?) a huge research arm in Reality Labs pushing the boundaries in AR, XR, wearable/portable technology. There have been plenty of public technology demonstrations showing all the stuff going on before the current AI trend.

Just over in Reality Labs - people learned how to build stuff from ASIC design up through optics, display, camera technology, highly miniaturized electronic design and power, all the fun stuff around 2D/3D audio and video stuff on wearable glasses (how do you provide head and world locked surfaces to your applications and not have it be so laggy that it gives people headaches?), making it all work over wifi, and then .. well imagine the lessons learned in what can and can't work in manufacturing the devices and where the pain points are in current technology. (And yes, a lot more I can't talk about, for hopefully obvious reasons.)

That knowledge is now embedded in a few hundred people, soon to be a few thousand people, who are being laid off. And that's just reality labs - the other people spread throughout the company and other technology companies have gathered a lot of experience about how to make and grow technology "stuff", what works and what doesn't.

At some point some of those people are going to make startups that do really amazing things. The technology underpinning a lot of what people have been trying to make now will get better (and I will argue that right now it is good enough - if you're willing to shift your focuses a little) and things will appear that will knock the socks off the current offerings. They're not what you would view as "founders" in the bay area tech scene. They're the people who know how to make things work, not necessarily the people who got rich early on in the scene.

Just like what happened after the dot com bust of 2001.

I've heard from a few people now that those they're hiring from Meta and other large technology companies are top notch and know what they're doing. You can't spend 10 years at a company that heavily invested in cutting edge R&D and not learn a thing or two.

I think its their loss, and .. eventually, everyone elses gain.

 

Sunday, February 15, 2026

Being accused of taking down a Data Centre, or "what happens when adrian is the only one who isn't oversubscribed on power?"

This happened circa what, 2008? 2007? I forget now. Anyway, it was a while ago now - and it wasn't my fault - but at the time it was pretty "wtf oh shit I'm in so much trouble" scary levels of scary.

So here we go.

Around that time I was running a little consulting / hosting company. I went into the data centre which hosted my equipment to install a second hand Dell I had acquired and tested. I plugged it into the rack, then plugged in the power, then pushed the on button.

Then bang. Then everything went dark. Then bang again.

Then I get a phone call on the VOIP phone in the data centre. It was their owners, asking me what the fuck I had done.

Anyway, it turned out that everything was dead. It wasn't just that the circuit breakers had tripped. A large part of their power equipment in the DC had also gone.

So yes, I was blamed for taking out a whole data centre. By plugging in one Dell server.

But why was I not in court over it? Why am I not still paying back the damages? Well, it turns out there's way, way more to the story.

First up they added a new rule - "You need to test equipment at this outlet/breaker before you install it." Cool, my server definitely passed that test. It worked just fine.

But then it turns out that although my rack was perfectly under the rated current limits, the other racks were not. Like, in any meaningful way. There were some other customers way, way over their allotted power. So when I plugged in my server - again, I'm way way under my own rack power allotment - I tripped some breaker on the distribution board.

Tripping that breaker meant that the other phase now took the brunt of the load. Again, my rack is fine, but everyone elses was apparently not, so it .. pulled very hard on that rail. And it immediately tripped the second phase breaker.

But that wasn't it.

Then, the second click was when they tried remote flipping the breakers back on.  The massive draw of power on phases when the computers in the data centre were powered back up caused some part of their power distribution setup to just plain fail. I forget the exact details here; I think the UPS was pulled on pretty hard too in that instant and I vaguely recall it also got cooked.

I had like, four? servers, a router and a switch. I was definitely not going to cause inrush problems. But the big hosting customers? Apparently they .. had more. Much, much more. Now this isn't my first rodeo when it comes to power sequencing of servers in a data centre - I had done this for like, a LOT of Sun T1s in circa 2000, staging how to turn them on a rack at a time upon a full power-off / power-on cycle event. But apparently this wasn't done by either the data centre or the hosting customers. All power, all on, all at once.

Now, I'm a small fry customer with one rack still paying the early adopter pricing. The companies in question had a lot more racks and were paying a lot more. So, this was all mostly swept under the rug, I stopped being blamed, and over the next few months we all got emails from the data centre telling us about the "new, very enforced power limits per rack, and we're going to keep an eye on it."

Anyway, fun times from ye olde past when I was doing dumb stuff but people were making much more money doing much dumber stuff at times.