Ah, the laughs (Xoogler since 2020).
It was a lot easier, at least last year: you'd use "flex" quota from your PA pool (product area) for Spanner and Borg, write some code for your server, a few configs here and there, and you'd be ready and serving.
About six years ago I had a resource manager deny me a database instance the very same day it became available for flex in another product area. I tried to "Hey Mister" resources from someone in that group to no avail. Eventually I wrote a high-durability key-value store on top of our source control system and told them they could give me my database or I'd be deploying to prod.
>Eventually I wrote a high-durability key-value store on top of our source control system and told them they could give me my database or I'd be deploying to prod.
That's diabolical. I knew Google should never have dropped the "do not be evil" thing
That video came out a few years before flex appeared I think at a time they were having a sort of “resource crunch” on the heels of growth spur following the GFC.
It wasn't that there was a crunch — that had always existed. There just wasn't all the tooling to implement anything like flex. At least this video was made after "buying Borg quota" was a normal thing. Before it, you had to "buy" regular machines and donate/assimilate them into Borg. Then after X days you'd receive your quota, minus a Borg "tax" of 10% to cover borglet and system daemons' overhead.
Hah, I think that at least at the beginning, the tools wanted to steamroll the Bigtable service, too. As the one that had to vet the changes every week, I had to go explain that it wasn't really neither our quota nor our usage.
For Colossus, sometimes large new clusters would go from very idle to super busy and under provisioned within days or few weeks, AFTER the steamrolling was already done, just because large users moved in. We then had to bring up new curators (masters) quickly, many times with quota that didn't exist. Quite often, when out of options, I ended up taking precious p360 quota from the storage monitoring user to mint production resources for Colossus. Fun times.
So before Borgmon existed, the quota unit was basically an entire machine? No virtualization?
Oh wait, this was probably 2008, VMWare had only just figured out JITed software virtualization a couple years prior. Makes a bit more sense now. And now the containerization thing makes a tad more sense: Google basically skipped over the "use QEMU" (or more recently, Firecracker) phase everyone else is now going through.
Borg was started in 2003 and by 2006-2007 was already the default way to run things, even though isolation wasn't perfect then. I think it started with chroot jails, then fake NUMA, then cgroups, which were written for it. It took years before all of web search moved to dedicated Borg machines (their quota belonged to you and nobody else could run on them) and, eventually, shared ones.
Ah, see, should've given 'em to us! Big rectangular state, you know, #3 machine owner at the time. We didn't charge overhead, probably because it never occurred to us to do it.
People brought machines, we gave 'em quota. Easy enough.
I know the service! The folks in NY had the state flag by their cubicle. My team in 2007-2011 was probably one of the largest users and I think I donated machines to y'all in the old Groningen cluster, before quotas were automated. GFS didn't charge for chunkserver overhead, either, and that mistake took years and lots of pain to fix...
It was easier to mint and carve out Colossus quota than e.g. Bigtable. I seem to remember that flex for Borg existed, but only in a few locations with enough capacity to back it. You couldn't just retrofit it in clusters where existing, large customers were already granted and using most of the quota.