# 364: AWS Billing Bug Sends Invoices to the Moon Duration: 64 minutes Speakers: Justin Brodley, Justin Date: 2026-07-31 ## Transcript [00:07] Justin Brodley: Welcome to The Cloud Pod, where the forecast is always cloudy. [00:10] Justin: We talk weekly about all things AWS, GCP, and Azure. We are your hosts, Justin, Jonathan, Ryan, and Matt. [00:18] Justin Brodley: Episode 364, recorded for July 21st, 2026. AWS billing bug sends invoices to the moon. Good evening, Matt. To the moon. To the moon. How you doing? [00:30] Justin: Yeah, enjoying my AWS bill that was very high for 3 seconds. [00:35] Justin Brodley: Yeah, no, I unfortunately here on the West Coast, I think they had fixed it by the time I woke up and actually looked at my bill, but I definitely would've had maybe a minor heart attack if I saw some of the numbers people were talking about, like billions of dollars or, you know, larger than a small, the GDP of a small country and billing errors. So yeah, you know, it's a, We don't really think much about those things until they, they surprise you. But yeah, that is actually our first story. So let's just jump into it then. So a bug in the AWS billing computation subsystem generated inflated billing estimates for some customers, with one Reddit user reporting a quoted estimate nearing $2.5 billion for a single month, while others saw figures ranging from millions to hundreds of millions. The issue started late Thursday, and an initial rollback attempt on Friday morning failed to resolve it, indicating the root cause was more complex than a simple recent configuration change. Amazon confirmed the billing estimates do not reflect actual usage or charges, meaning affected customers will not be responsible for the inflated amounts shown in the console. Amazon has not disclosed whether any accounts were suspended or paused due to the billing errors, leaving open questions about operational impact during the incident. The event highlights the importance of billing system reliability for cloud providers. Yeah, I mean, so typically Amazon doesn't bill you in the middle of the month, so it's a pretty low risk that you were going to get billed or invoiced directly on that date. Unless you happen to already be overdue on a payment and you were having to update your credit card at the same time. I don't think that's really a big risk for this particular scenario. [01:57] Justin: But even then, they wouldn't have charged you for it. [01:59] Justin Brodley: Yeah, I don't think they would have charged it, and your credit card would have said a billion dollars, you're silly. So you don't have that, you don't have that kind of credit limit. [02:07] Justin: Well, let's have it pulled from your bank account. [02:09] Justin Brodley: Well, I mean, I, I don't know who— I mean, other than Elon Musk, who has that kind of money in their bank account? [02:14] Justin: Well, Well, that's the problem. Just overdraw. [02:16] Justin Brodley: Yeah, just overdraw. So the big penalty from your bank. I mean, the reality is that we don't really think about the fact that Amazon makes mistakes in billing, which they do. And I've, I've heard stories from people, it's never happened to me personally. I'm kind of sad. I'm hoping for a bank error in my favor someday. But, um, you know, I've, I've seen several times where, you know, customers have told me that, oh yeah, they misbilled us for something. But typically it's something around like, oh, well we had a special contracting price that we negotiated that was put into our contract that someone didn't properly document and then they updated something and that's what broke it. But you know, they, they, I've been stressed before that if you have a lot of special pricing, you should definitely be auditing it 'cause it can break. [02:54] Justin: Yeah, I've seen that and I've seen a bunch of like automated, like S3 pricing, here's 3 cents back, you know, and or $5 back, back when, uh, we ran our own reseller program and we'd get, have to process the bills from Amazon after the fact. So I've definitely seen credits in that way, but the bigger question is how do you actually validate your bill? You know, and that's kind of what I think this made me think about is how do I actually validate that what Amazon's saying is correct is correct for my small AWS account? Yeah, sure. I run a spot instance or whatever else, so I don't really care. But, you know, on a thousands of resource, you know, level account, How do you validate that, you know, the EC2 instance is the right price at the right, at the right level, you know, at the right instance type or the Load Balancer and the, how do you level manage your LCUs and know how many LCUs you're using? So you get into this whole other world of how do you audit your provider? 'Cause you know, if you go to the mechanic, you at least can say, oh yeah, yeah, they did replace my oil. You might not be able to audit down to the number of quarts they put in, but you know, you can audit what they've done versus Amazon. And with millions of line items in there, you can't really audit as well. [04:07] Justin Brodley: Yeah, it definitely becomes a problem at scale. And also, if you're thinking like, well, I have a tool like Aptio or Code Zero or Cloud Zero, whatever one of those other tools, like they also typically are using the CUDs or the CUR reports. And so they typically aren't recalculating what your bill should actually be. And so I know bigger customers have built,, you know, kind of their own mini billing system to kind of keep track of their own costs, especially with AI and some of those things. I could see this becoming something more and more people want to do because tracking tokens can also be an area ripe for, you know, challenges of where, well, this query yesterday only took this many tokens and now today all of a sudden the token count changed and like, you know, being able to narrow that down, hundreds of thousands of transactions against, uh, AI cache, uh, could be, save you a lot of money. So I imagine this will become something maybe people are more interested in the future too. [04:57] Justin: Yeah, be curious to see how this evolves over time as people kind of figure out how to, how to audit the auditor, essentially. [05:05] Justin Brodley: Mm-hmm. All right, Henrico County, Virginia might not be a place that you know a lot about, but I bet your data lives there if you're in the US East 1 Virginia, 'cause there's 37 data centers apparently in Henrico County and with 17 more planned. So I can only think that this is, must be, uh, the county where most of the major cloud providers, or if not Amazon, at least your cloud provider may be there. They're apparently asking county employees and schools though to conserve electricity after a 25% rate increase, which is set to begin July 1st, adding an estimated $5 million in costs for the next fiscal year. The situation highlights a direct tension between the data center growth and local Azure costs, with residents and government facilities absorbing higher utility rates, likely tied to the power demands of the nearby facilities. The proposed expansion includes converting Civil War battlefield land into data center space, raising questions of the land use and community pushback in addition to energy concerns., and this is just one more example of the growing backlash against AI data centers and data centers in general happening across the United States between the noise pollution, the light pollution, and usage of water, especially in areas that are drought-stricken like my lovely California. You know, it's becoming a bigger and bigger topic. And now you're telling, you know, a bunch of kids they can't use power in their classroom where they're supposed to be learning because they didn't expect a 25% price increase, uh, is pretty Pretty crappy scenario. [06:23] Justin: I mean, I think it becomes a larger conversation too of not just, you know, the schools, but how's it affecting the people in the area, you know, the people have lived there for years that all of a sudden now they can't afford their power bill, you know, and then everything else is like, okay, so you're now not just affecting the people, it's what is the actual effect of the data center? It's, you know, the actual building of the building, initial water load, let's say it's a fully, you know, one of those new data centers that you only load, you know, the water once every 20 years. So they say, you know, the one that when they did in Georgia, everyone complained in the area that their water pressure was low for 6 months and then it went back to normal once they loaded it. It's like, there's all these other effects in the area that I think you have to think about. But I also think you can't be one of those people that says, I want to use AI and I want to use AWS and everything else and just say, don't, no, but don't do it in my backyard, you know. So there's a middle ground. 37 data centers in one county feels like a lot, I'm just saying. So like, there's somewhere I feel like that one's a little bit more skewed over there. [07:28] Justin Brodley: Oh yeah, I mean, Henrico County, I think before data centers arrived, was, you know, not a huge, like, highly populated area, which is why it was chosen. And so now it's been populated by a lot of data centers, but there are still people who live there and, and do different things in that space. [07:43] Justin: So it's Well, Henrico County looks like it includes Richmond, or is it just outside Richmond? [07:49] Justin Brodley: It's just north of Richmond, as my understanding. [07:52] Justin: It looks like it kind of loops around Richmond. [07:54] Justin Brodley: Yeah. [07:54] Justin: One of those weird ones. [07:56] Justin Brodley: Yeah. But you can see a lot of it's just green fields and, you know, it's, it's empty area for the most part. Yeah. A lot of the county. So, but yeah, you definitely zoom in and see a lot of data center facilities if you look, uh, know what you're looking for. Anyways. [08:12] Justin: Yeah. [08:12] Justin Brodley: Uh, don't, don't take power away from school kids. And I, I do like the, you know, at least it sounds like in a lot of the newer municipalities where they're agreeing to put these data centers in, they're saying, you know, we're not pushing rate increases down onto the, the general population that, you know, if rate increases required because you're using so much power, you are gonna pay for it, which I think is the right way to handle that. [08:32] Justin: I do too. [08:34] Justin Brodley: So, you know, I definitely, you know, and that's why I think you're also gonna see probably the next 10 years a lot of investment in, um,, you know, 20 to 30 megawatt mini nuclear plants being put into these data centers too, which will have a whole different set of, uh, political ramifications for people who are gonna be pissed off about that. So it's, we're definitely just the beginning of the AI data center pushback movement, I think. [08:57] Justin: Well, we talked about for a while the, was it small nuclear data centers? They think they were even smaller than the fi— that in the '30s they were like 5s or whatnot. [09:05] Justin Brodley: Yeah. And they're, they go from 5 to 10, you know, but like typically a data center you might need, you know, especially with the GPUs that are so powerful. Hungry, you might need a 10 or 15 megawatt setup and then you're gonna have to redundancy. So now you're talking about how two of them, so you're having this thing on 30 megawatts of capacity in that space. So, you know, again, I, I don't know, I don't know about those technologies. I know they're coming, but, uh, I assume, you know, even though they are much, much safer than traditional nuclear power, people are going to hear nuclear power and freak out. [09:33] Justin: Freak out. Yeah. Yeah. No one wants to hear a whole conversation about stuff. [09:37] Justin Brodley: Exactly. They just wanna hear, I know what nuclear is bad. So there you go. Well, uh, moving on to AI is how ML makes money. Uh, OpenAI has trained a new GPT called GPT-RED as an internal-only automated red teaming model used to find prompt injection vulnerabilities and generate adversarial training data at the compute scale. Some of the largest post-training runs. GPT-RED uses self-play reinforcement learning against a population of defender LLMs with GPT-RED rewarded for successful attacks and defenders rewarded for resisting them. Forcing progressively stronger and more diverse attack discovery. Incorporating GPT-RED into training produced GPT-5.6 SOL, which shows 6x fewer failures on the hardest direct prompt injection benchmark versus the production model from 4 months earlier. And it fails only on half a percent. [10:24] Justin: That's correct. No. [10:26] Justin Brodley: 0.0005% of the GPT-RED direct prompt injection attempts. Uh, in generalization tests, GPT-RED achieved an 84% attack success rate on novel indirect prompt injection scenarios against GPT-5.1 compared to 30% for human red teamers on the same task. In a real-world task against an AI-powered vending machine agent similar to Anthropic's Project Bend, GPT-RED successfully changed item pricing, created a fraudulent listing, and canceled another customer's order with the vulnerabilities disclosed and safeguards now being tested. So, good to see these tools. I assume this maybe comes— turns into something like Mythos for OpenAI at some point in the future to make this available to anybody. But, uh, you know, the ability to attack and attack from multiple vectors and chain attacks is only increasing dramatically at this point. [11:10] Justin: Yeah. And do you think we're going to talk about later too? It's just the number of vulnerabilities out there in the world are just increasing more and more. And I think a lot of it's just these tools are getting used and you can run more and more automated testing against yourself and specifically the daisy chaining stuff, you know, 4 lows now can really very much equal a high or critical or your system getting completely owned, 'cause you're able to daisy chain those things in. So you get in with a low that gives you some level of access, then you use a medium to high to escalate your privileges, and now you have officially owned that system. So, you know, leveraging, letting AI do this, and it'll be interesting to see how Mythos and, GPT-Red kind of compare later on, but I don't think most of us are going to be able to play with those models for a long time in a real way. [12:01] Justin Brodley: Now, unless you, uh, are approved by the government, so yeah, right. OpenAI has also released its first branded hardware, the $230 Codex Micro, a collaboration with WorkLadder built on their existing Creator Micro keyboard line rather than a fully in-house design. The keyboard's key feature is 6 frosted color-coded keys that provide status updates on up to 6 concurrent Codex agent threads. White for idle, blue for processing, green for completed, and amber for needing human input, and red for errors. 6 additional programmable buttons handle common Codex actions like accepting or rejecting changes and branching threads, plus a push-to-talk button for audio prompts. Users can remap these and access 5 additional customizable layers for general shortcuts via 32 included keycaps. The device addresses a workflow problem for developers running multiple AI coding agents simultaneously, offering a glance At-a-glance monitoring as alternative to keeping several browser tabs or a laptop open to track agent status. This is a desktop-focused accessory that complements the router than replaces mobile monitoring options like the ChatGPT app. I don't get it. [13:00] Justin: I have a Stream Deck from 4 years ago that I think is more powerful than this thing. Like, I'm pretty sure I could wire up this to Claude or Codex and do the same thing, but mine are little LCD screens. I don't know what they are actually. OLED screens probably that I can tweak and set to whatever I want. Now I don't actually use the device that much, but it's plugged into my, plugged into my computer via my USB-C hub and it's there. [13:27] Justin Brodley: And I mean, and that costs less than this device, does it not? It has LCD screens for— [13:32] Justin: yeah. I mean, mine's an old one. [13:35] Justin Brodley: Yeah. I mean, I think a brand new one with 9 buttons was like $140. So yeah. [13:40] Justin: And I have 15 on mine. So a 15-key one is $119 at Best Buy. [13:47] Justin Brodley: Yeah. So that's not that bad. [13:48] Justin: Yeah. [13:49] Justin Brodley: I guess I don't understand, like, you know, the, the way they describe this is like, well, if you have multiple coding agents, you can see at a glance. I'm like, well, it only has one color. It doesn't have like different sections for different colors. Or I guess some of these buttons you may be able to code different colors and tie those to different agents, but only like part of them, which is weird, but only a part of them. Right. And then, okay. But like. How hard is it for me to press X on my keyboard or Y on my keyboard or N on my keyboard versus needing this? I just, I don't understand the point of this. And I, like, I wanna support the idea of like innovative hardware devices. I just don't think this is— [14:23] Justin: I don't think this is innovative. [14:25] Justin Brodley: No, I don't think it's innovative at all. [14:26] Justin: They don't do something. [14:27] Justin Brodley: I mean, and they didn't even, they didn't even make it. It's a partnership with Work Louder. [14:30] Justin: So, and that's after they're getting sued by Apple, which I think we decided not to talk about last week. Yeah, we didn't talk about that, but like they're getting sued by Apple for stealing their people. So like, is there something else coming? Is this just an initial thing that maybe they are not using their internal team for, hence the partnership, which is the only thing I could think of, I guess. [14:50] Justin Brodley: I mean, I, I, I had to see, like there is on the WorkLadder website, there's a picture showing the, the different buttons lit up in different colors between idle, thinking, complete, and needs input. But even then, I'm like, I, I'd have to see someone like really using this in their workflow to understand how it's really improving anything. 'Cause I can only read so fast and I don't wanna just blindly approve what the agents are doing. So I, again, I, I don't think it's innovative. I don't think it's cool. I'm waiting for them to come up with something better. So nice try, strike and a miss for me personally. [15:20] Justin: Yeah, I don't see the value of this. And now I'm actually more curious to see if I can get my Stream Deck to link up to Claude Code. So you've kind of gone the other way. [15:29] Justin Brodley: Yep. Well, uh, we have two articles here, uh, about Moonshot AI's upcoming KIMI-K3 model. Uh, apparently it's reported to perform on par with or exceed Anthropic's Opus 4.8, according to sources cited by the Financial Times. It's expected to be the largest open-weight AI model out of China, with parameters ranging between 2 and 3 trillion. Predecessor KIMI-K2 already ranks competitively on open-source benchmarks, and K3 aims to further narrow the performance gap with closed-source frontier models from OpenAI and Anthropic. Anthropic. Moonshot is reportedly raising a new funding round at $3 to $1.5 billion valuation, up from the $20 billion in May when it raised $2 billion, reflecting continued investor interest in open-source AI development. This release comes as enterprise leaders debate the cost and data privacy trade-offs of closed-source AI subscriptions, with some executives recommending open-source alternatives like Moonshot, Deepseek, or Z.AI for organizations wanting to train and control their own models. And then K3 Tech Blog came out with a bunch of their metrics and benchmarks about how they think it's going to perform. Dockerform, and they are definitely targeting trying to hit Opus 4.8. Um, now I can't try these, uh, it's not available to me in any place that I can get it, and Kyma actually shut down signups to their cloud because the demand was so high. But definitely the overall stock market has not been favorable to this this week because again there's a lot of companies investing a lot of capital, and so these cheaper models put that business model at risk, and so the, uh, the market is appropriately reacting this week, but I'm definitely excited to get my hands on KIIME K3. If it's as good as Opus 4.8, which is pretty darn impressive, I'll be maybe using that exclusively. I don't know. [17:03] Justin: Yeah. I mean, the price point, if I remember seeing it, was much cheaper. [17:07] Justin Brodley: Fractions of, of the cost of Opus. [17:09] Justin: So yeah, I was just gonna say like 25%, but I'll, I'll not throw a number out there cuz I have no memory actually at this point of it. But yeah, no, we have in our notes, we have 30 cents per million tokens input cache. Then again, figuring out how pricing and everything else works of cloud, of, uh, not just cloud, but of, uh, AI is also hard. [17:29] Justin Brodley: Yeah. So input tokens, which are cache-missed tokens, uh, KEMI 2.7 is, uh, 95 cents per million, where Opus 4.8 is $5 per million. Output tokens for KEMI 2.7 is $4 per million, while Opus 4.8 is $25 per million. The context window, uh, for Kemi is much smaller though. It's only 200, 262 tokens on the Kemi 2.7. I don't know if they mentioned context increases for Kemi 3.0, but you know, for those pricing differences, I can unload and reload context multiple times before I care. [18:02] Justin: Yeah. I mean, it's gonna be a no-brainer to switch at that point. Yeah. So I feel like the rubber's gonna be where it meets the road on July 27th to actually see how it actually is on these platforms. [18:15] Justin Brodley: How it actually works. Yeah. And you expect to see a lot of FUD coming out about this and, you know, both Anthropic and OpenAI, I'm sure are going to be pushing. Oh yes, they do plan to have a 1 million context token window for KIMI-K3. So. [18:29] Justin: But I still don't want a million. I've done something wrong when I've got a million. [18:35] Justin Brodley: I don't know. I agree and I disagree. So like I do, I typically clear context a lot, but there are times where If I'm doing a lot of debugging, it's nice to be able to just not worry about parsing it down to just the, the tiniest bits. Like here, just a, here's a console dump. Just take it, figure out what's wrong with this thing. There are advantages to it as well as in long coding sessions, the longer context, if you keep it clean and maintained is actually pretty advantageous. Like, you know, if I can go jump between agents that have been running for a while, then, you know, I got a code rabbit finding back and then I go back and look at it. I'm like, oh, okay, well I don't have to reload context now. So there are some advantages to it. I agree with you. Most things you do should be in the 250 limit, but when you do need it, it's nice to have. [19:14] Justin: It's nice to have. It's almost one of those things I would like it to say, hey, you've reached 200,000 or you've reached 150,000. Do you want to keep going? You know, or should I compact now? You know, and take it that way versus just waiting for it to keep growing and growing and growing. Because a couple times I've caught it. Pretty large and I'm like, hold on, I've just been using this chat because I've been doing 17 different random things in my head. Let me stop and actually go properly fork and do different threads and go from there. [19:46] Justin Brodley: Yeah. Well, I, I, yeah, agree to disagree, but I don't mind it that much, but I get what you're saying, Matt. So, all right. And then finally, OpenAI is launching the ChatGPT for Small Business program, bundling virtual training webinars, in-person AI academies, guides, and curated partner integrations from Dropbox, Microsoft, Shopify, Intuit, Slack, Atlassian, and Wix. ChatGPT Work, OpenAI's multi-step task agent like Cowork, is now available to small businesses and runs on GPT-5.6, positioning as OpenAI's most advanced model available across all business subscription tiers. From last year's Small Business AI Jam events, OpenAI reports 78% of participants built a functional AI workflow in a single day, and 42% saved more than 5 hours per week using AI tools. Use cases highlighted include converting voice notes into Slack messages, generating real-time market competitor tracking sites, evaluating inventory for product or marketing ideas, and building training presentations from customer review data. The program targets a segment often lacking dedicated IT or automation resources, framing Agentic AI as a way to offload tasks like marketing, accounting, and operations. [20:48] Justin: You know, I work with a lot of small businesses, and a lot of small businesses are trying to leverage AI to do specifically those things and just the general interim stuff that a large company has hundreds of people in different departments doing it. Where when you're a 5, 10, 1, you know, 30-person company, you're in those ranges, you don't have full departments, you don't have people for that. So I think it's great that we're trying to teach these people. I think you're gonna get a lot of pretty innovative things out of it 'cause you're gonna get a small business that just doesn't have a marketing department or doesn't have one of these things that has figured out how to leverage AI to build a, you know, a proper workflow and go from there. [21:30] Justin Brodley: Yep. All right, let's move on to security. Windows zero-day drops the same day Microsoft releases a record number of patches. They released 570 security patches, which is a record for a single update cycle. And a zero-day exploit surfaced the same day affecting the Windows User Profile Service. The exploit called HiveLegacy allows a low-privilege account to modify an administrator account's class registry hives, which controls file associations and behaviors in Windows Explorer. Exploitation requires attacker to know credentials for one account and the username of a second account on the same machine, limiting but not fully eliminating the practical risk. I mean, in general, 570 security patches is a lot of security patches. And I can definitely thank AI for all of those patches because this is raising the priority on a lot of low and what would previously have been considered a not risky vulnerability are now being chained into bigger attacks, which we talked about earlier. And now those are actually get patched much, much faster. But, um, I also heard lots of complaints from people about 570 security patches blew up a lot of like indexes and security patching tools that had never had such large amounts of data to process at one time. So yeah, it was a bit of a rough week for the cloud, you know, Windows patching, uh, ecosystem. [22:41] Justin: Yeah. I mean, the statistic I used to hear was that from the time the vulnerability was found or announced until the time it was attacked was like a couple days. And I think most recently I heard, the last time I heard, it's like under 20 hours. So essentially from the time they announced this until then, you had 20 hours to patch it. And most companies cannot patch. I mean, I would go as far as saying all companies, but, you know, maybe there's a few here or there they can, cannot patch that quickly, let alone get the Vault, get the updates from Microsoft and everything else. So you really have to be focusing on— [23:20] Justin Brodley: let me introduce you to a little tool called Microsoft SQL Server. And, uh, the problem with that tool is patching it requires downtime typically, and that's an even riskier problem. [23:31] Justin: So yeah, so you got those problems. I mean, hopefully your SQL Server is not public-facing, as you've talked about in the past, and it's behind multiple layers or multiple firewalls. So somebody has to be in your system, but that's where you get the daisy chaining stuff where all these loads between systems slowly become a harder and harder issue to deal with. [23:51] Justin Brodley: Yeah, I— it's only get worse if it gets better. But I definitely, you know, if you have been sort of lackadaisical about patching at your organization in the past, I think, you know, that, that era is now ending. And I think, you know, I don't know if we need to get to the level of FedRAMP level craziness, but I— something between where most companies are at and where FedRAMP requires for patching, I think is going to be a requirement for most people going forward. [24:15] Justin: I mean, most contracts that I saw were 7 days for criticals, 30s for high. I feel like that's what most enterprise customers are requiring, and it wouldn't surprise me if that shrinks more and more over the time, over the next like 6 months. [24:30] Justin Brodley: Uh, curious to see if it is a momentary blip where we have a lot and then it kind of dies off, or if it continues to be a high volume. But I mean, the reality is Product like Windows, I mean, it's decades old at this point. It's big. It's massive. Has a huge footprint. And the, the reality of that is that you are always going to have more risk in that product, which is why a lot of, you know, there's a lot of interest in like rewriting kernels into Rust, into different technologies and C++, and 'cause it helps reduce some of those risk areas. And that's why you see a lot of projects converting to Rust into other programming languages that are more type safe, more memory safe. Because of the complexity of these things. So it's an interesting time. I think there's a lot of refactoring happening right now with AI. AI is exposing a lot more. We're all responding to it a lot more. And either this is a temporary bubble or it's of the future that we'll all be living in forever. [25:18] Justin: Didn't Microsoft announce that they're rewriting their, their kernel to Rust, or was it just they were adding Rust components? [25:24] Justin Brodley: Uh, I thought they were just adding— I know Linux has been moving to Rust. [25:28] Justin: Why is Microsoft rewriting Linux? I think they're adding, and I think the goal is to eventually fully convert. [25:34] Justin Brodley: The Linux community is the one who's, you know, moving major portions of the kernel to Rust. Uh, Microsoft is not rewriting. [25:41] Justin: Microsoft to replace all C and C++ code with Rust by 2030. [25:46] Justin Brodley: Got it. [25:47] Justin: Including the kernel, NT/Windows kernel. [25:50] Justin Brodley: They're saying that it's not everything, just the highly vulnerable security critical parts of the kernel and core libraries with Rust. At least, at least from this Gemini summary of my search. And then the 2030 goals there as well. But I don't, yeah, I don't, again, Rust is a much more stable choice than C++, so I get it. Or C even, which I'm sure a lot of the kernel's written in C natively. [26:14] Justin: Probably some assembly just to screw with people in there. Highly efficient. [26:18] Justin Brodley: Yeah, super highly efficient. All right, moving on to cloud tools. Uh, 1Password for Claude is available, which lets the AI agent complete browser logins and tasks without ever seeing the actual password for one-time passcode. Events are injected directly into the page at runtime while 1Password remains the source of truth. Uh, and I will tell you, that's all I'm going to say about this. Uh, it exists. I can't make it work. I've tried. I followed their documentation. I have the plugins. I made sure all of them are updated. I made sure I had the latest 1Password. I made sure I had the MCP turned on. I configured the MCP and different things. Like, I cannot get this to work. [26:52] Justin: You probably spent more time on it. I think I spent all of, uh, 15 minutes, got annoyed by it. And said, this doesn't work. I'll try again when somebody has a better post or they've updated or I reboot my laptop or something. [27:04] Justin Brodley: It was partially annoying too because, you know, they did this blog post about 1Password for Claude and then none of their documentation got updated. So none of the examples of how to do it are for anything other than OpenAI Codex and Quro. And I'm like, but you just announced this. Why did you not update the documentation to include the Claude steps? [27:21] Justin: Different pizza box team, Justin. [27:23] Justin Brodley: I, I know, but like they're saying it's all you need is the desktop app, the browser extension, the cloud desktop app, and the cloud in Chrome. And I can tell you I have all 4 of those and this does not work. So try, try harder. I even watched the video. That's how much I cared. [27:38] Justin: You went further than me. I got annoyed and said, I gotta leave. So screw it. I hope this works and it doesn't work. So I'll move on in life. [27:45] Justin Brodley: Yeah. And then HashiCorp is launching TF Policy in public beta on HCP Terraform, which is a declarative policy as code framework using HCL. For those of you who are going like, but they have Sentinel, you're correct. This is an alternative to Sentinel, or OPA, letting platform teams write governance rules in the same syntax used for infrastructure definitions. Key new capability is relationship-aware policy evaluation, allowing rules to span multiple connected resources. For example, requiring every IAM role to have at least one attached policy rather than checking resources in isolation. The framework supports data source lookups during policy evaluation, so policies can reference external context like improved AMI lists or organizational inventories rather than relying solely on what's defined in the Terraform configuration, which that's a nice feature. [28:30] Justin: That's a great feature because that was a major pain point. [28:33] Justin Brodley: It was. TF policy adds controls to block unapproved provider and module downloads before use, addressing supply chain risk by enforcing that dependencies come from approved private registries. And policies can be evaluated post-deployment, checking provider computed values like generated ARNs against organizational standards. Which addresses gaps where plan time checks alone are insufficient. HashiCorp is providing an agent skill on GitHub to help teams author and test TF policy files or migrate existing Sentinel policies, easing adoption for current GCP Terraform customers. So I suspect that Sentinel is going to go away or at least be heavily deprecated in favor of this method. [29:09] Justin: I feel like there hasn't been a ton of improvements to Sentinel in a long time. [29:12] Justin Brodley: Well, I know it is one of my biggest complaints about it is that the language is just annoying comparatively to HCL. So like, this is so much better from that perspective. [29:23] Justin: Well, you're expecting somebody to write Terraform or a security person that probably has maybe, or hopefully has some knowledge of Terraform, that to learn a completely other language for this one specific tool is only useful in so many ways. So like, like, I feel like the people I realistically saw using it were the centralized public cloud COE or like any of those types of teams. Where they were able to do it, but then they were used to Terraform, managing Terraform, and here's another language just to screw with you. And like, it just never was clean in the ecosystem. [29:56] Justin Brodley: Yeah, I agree. And it always felt like a bolt-on. I mean, I think that was the biggest problem with it was like they, they needed to create a solution. Someone created it quickly and then they embedded it into Terraform Enterprise and said, there you go. And then everyone was kind of like, okay, but OPA is better, or You know, whatever the, the comparative to Amazon has a language for this. I'm forgetting the name of, that's kind of also policy as code type that people also heavily adopted. So I, I just think Sentinel just didn't win. And so by making this more closer to Terraform, I think they have a better chance of getting people to actually adopt it. [30:29] Justin: I'm trying to think what you're thinking of for Amazon policy as code. Like they have like SCPs, but it's a mathematical proof. [30:40] Justin Brodley: Concept they had. [30:41] Justin: I remember, I didn't think it was a project. I thought that was like some backend thing. [30:45] Justin Brodley: It's, it's part of it. [30:47] Justin: Cedar. [30:47] Justin Brodley: Cedar. Cedar policy. Yep. Cedar policies. That's what it is. [30:51] Justin: I forgot about that. [30:52] Justin Brodley: It's still out there and there's, there's many companies are, are very happily applying Cedar. It also has a lot of value in some of the Gentec use cases too, cuz one of the advantages of Cedar was that it, it doesn't just look at user and you know, your password as an identity. It also can look at like the posture of your laptop and a bunch of other things that can now factor into your policies for AI. So there's some advantages to it. All right, let's move on to AWS. AWS Lambda announces self-managed code storage. Lambda now lets functions and layers reference code directly from customer-owned S3 buckets instead of copying deployment packages into Lambda-managed storage, removing the 75 gigabit per region storage cap for those using this mode. Teams with many functions or larger players no longer need to file support tickets to raise storage quotas. Skipping the internal copy step also reduces the function activation time after creates and updates, which benefits customers with large deployment packages or frequent deployment cycles. So it requires setting S3 object storage mode to reference via CLI, CloudFormation, SAML, or SDKs, plus granting the Lambda service principal S3GetObject and S3GetObjectVersion permissions on the source bucket. Console-based updates are also supported for existing functions. No additional Lambda fees apply for self-managed storage. Customers pay standard S3 storage rates and cross-region data transfer costs where applicable. AWS has also raised the default Lambda managed code storage limit from 75 gigabytes to 300 gigabytes per region per account, benefiting customers who don't want to do this self-managed storage option, which I don't know who you are, but just use your own buckets, which I was always annoyed that you couldn't use your own buckets. [32:24] Justin: No, I would always upload the zip somewhere and then it would take the zip and move it over somewhere else, which I guess was the Lambda managed, which I never thought about. Though I've never hit the 75 gigabyte limit. [32:35] Justin Brodley: And that's a lot of Lambda code. I mean, that's— [32:38] Justin: well, I've done a lot of development where I used to like, when Terraform first came out, I would like have my Lambda built into my Terraform. I would just do a terraform apply every time, which would just zip up the folder and shove it into Lambda. So I've never hit that limit, but I guess there's people that do it. [32:55] Justin Brodley: I guess if you're fully, fully Lambda-based, if you're fully Lambda-based and you're, you're moving around large objects, I could see you hitting that limitation. But, uh, yeah, I, I just always hated that whole idea of like, I know it's, you're storing it someplace I can't see, and now here's a bucket you can use instead, which also makes it simpler for me 'cause now I control the bucket, I control the encryption keys, I can control everything I care about. [33:17] Justin: I assume it's also because they don't want to have, which, you know, they're gonna take a, Hey, my Lambda stopped working. Well, you deleted the bucket that had your Lambda in it. You know that that's what they were solving for. [33:29] Justin Brodley: I mean, that's, but that's okay. Like, don't bat on me, fuck gun to myself. But I mean, it's still nicer when you have full control in my opinion. [33:38] Justin: I'm not disagreeing, but I think we're, we're not the people that are going to open a support ticket for that. [33:43] Justin Brodley: True, true, true. All right. Amazon MQ now lets customers configure EBS storage size independently of instance type for RabbitMQ brokers, addressing longstanding limitation where storage was tied to computing size. And I only mentioned this story, uh, because we talked about RabbitMQ last week and we said, uh, yeah, you know, they have that RabbitMQ and I say, I said, yeah, I don't care about that. And apparently, uh, Amazon still does enough care and cares enough to still build features for it. So there you go. You know, get different storage tiers. [34:11] Justin: And it's just a quality of life that you could actually define it where I guess before you couldn't. Never ran into the issue, but I'll take it. [34:18] Justin Brodley: Yeah, I mean, I think it was dynamically managed based on the volume and the number of messages in the broker. And so, you know, you couldn't scale up automatically, you know, ahead of a, like a, if you knew there was a large capacity thing coming up, it's very similar to Redshift and how Redshift clusters worked or Elastic nodes as well. And so I think everyone's been sort of asking for that for those in those services. One time is like, I don't want to be tied to your storage tiers. I need to be able to use more or use less and have better cost transparency. So I, I wouldn't be surprised if you didn't see this following other services that sort of always have these coupled storage/compute models where Amazon at least feels comfortable that they can manage the risk. But yeah, glad to see it. Amazon CloudWatch Logs now automatically tiers data into standard and frequent access and archive instance access based on usage, removing the need to manually filter or export logs to cheaper storage elsewhere. Data shifts to infrequent access after 30 days without access and to archive instant access after 90 days without automatic promotion, with automatic promotion back to standard for 30 days when older logs are queried again. Query experience remains consistent across all tiers, letting teams keep both high-volume logs in CloudWatch long-term, leveraging tools or maintaining separate storage systems. So yeah, thanks. Appreciate that the thing that we knew it was always running on top of, which is S3, that had this capability that you wouldn't let us use, now lets us use it. So appreciate that. [35:44] Justin: That's a really long clap. [35:46] Justin Brodley: It is a long clap. [35:48] Justin: That's my official response. Really? It took this long to do that? [35:53] Justin Brodley: Yeah. Well, keep that clapper nearby, Matt, because Amazon Cognito now allows password hashes to be included in CSV user imports, letting migrated users sign in immediately without existing credentials instead of being forced into password reset on first login. Supported hashing algorithms include bcrypt, scrypt, Argon2-ID, and PBKDF2. With SHA-256, covering most common formats used by legacy identity systems and custom authentications. And so finally, I no longer have to now plan lengthy communication to all of my customers because I'm trying to move off of this platform to Cognito and I need you to change your password because I can't just import your existing password. Thank you. Thank you very much. [36:38] Justin: Okay. So if you, I can cut it off, but it doesn't cleanly, cleanly end the file. [36:43] Justin Brodley: Yeah. [36:43] Justin: Yeah. [36:44] Justin Brodley: I'll have to find a shorter clip, but yeah, this is, this is like, thank God. I, this was such an annoyance. [36:50] Justin: Can you export though from Cognito? Like, cause I feel like a lot of people start with Cognito to kind of set up a basic PRC and get things moving and then they want to move to Okta or re-move over. I've never thought about it in that way. I know there was. I think there was a way, 'cause that was kind of the hacky way to back up and restore in a DR region, but it wasn't ever good. [37:12] Justin Brodley: No, you cannot export hashed passwords from Cognito. So you would, if you're going the other way, you still have to do the painful dance. But again, I'm not, that doesn't shock me, but at least, you know, if you want me to move off of Okta to Cognito, telling me to re-have all my users change their passwords is just a terrible scenario. So how many of those emails have you seen over the years where, oh, we're changing our authentication system, you need to hit the reset password button to reset your password? Uh, I, I don't know if those people are probably what those people are running from. Cognito is my assumption on those, but yeah, maybe they're running to it and I just said no. [37:44] Justin: So I just also always assume that those were, oh, we got breached, so we got our password file and, uh, good luck. [37:49] Justin Brodley: Well, that's my other assumption too. It's like, oh yeah, it's one of the two scenarios. And I'll be getting a notice about my free credit monitoring any day. [37:58] Justin: I know I'm waiting for one. I think my last one just ran out, so, you know, I'm waiting for I'm probably missing one that I should go back and find. [38:05] Justin Brodley: Yeah, probably. I mean, it seems like every other day I get a dark mode, a dark web monitoring about my email address. So it's just one of those things. AWS Control Tower Account Factory for Terraform now automatically reapplies account customizations when accounts move between organizational units, eliminating the manual retriggering step that previously created operational overhead and configuration drift risk. Enable the feature by setting aft-customization-triggers=account-move in your AFT configuration., and the reapplication process skips bootstrap and provisioning phrases, running only global and account-level customizations for faster execution. Uh, yes, thank you. This was dumb. Another, another great clap-worthy, uh, announcement for sure. I didn't mean to actually do the clapping. I just said it was clap-worthy. [38:52] Justin: I just, uh, I had to have fun. Oh, here you go. Sorry, I should have done this one. [39:00] Justin Brodley: It's a much bigger applause. It makes more sense. [39:02] Justin: Yeah. [39:02] Justin Brodley: So yeah, that's good. [39:04] Justin: No, this is one of those things that I was shocked they didn't have when they announced it. I was like, really? They didn't do this? But I guess I've never played fully with Terraform Account Factory. [39:13] Justin Brodley: I went through a process where we were moving Amazon accounts around an organizational structure to get ready for a divestment activity at a much prior job to this one. And this was a problem. It's like, oh, you decoupled it, and now all of a sudden none of our rules are being applied properly. And we only found that because some developer figured it out first and applied, obviously, deployed stuff he shouldn't have. So it was like, oh, that's fun. So anyways, well, if you, uh, heard her talking about water or electricity usage earlier and you were also thinking about all that pesky water that, uh, Matthew mentioned, you know, low water pressure as they filled, uh, water cooling systems, AWS Sustainability is now letting you actually track water withdrawal data. So if you have customers or people who are upset about your AI usage is Feeling the water. Amazon Sustainability now adds water withdrawals data alongside existing carbon emission metrics, giving you a fuller picture of the environmental footprint tied to your workload. Data is broken down by AWS region, service, and account, and it's reported annually through both the console and the API, allowing you to integrate into existing reporting workflows and dashboards. Feature is free in all regions where AWS Sustainability is available, removing cost as a barrier to loading adoption of ESG and sustainability reporting teams. Lower withdrawal volumes reflect data center efficiency improvements, giving customers a way to track AWS infrastructure efficiency gains over time as part of their own sustainability disclosures. [40:33] Justin: I think it's going to be interesting where you're kind of going to be able to reverse engineer what, what and how Amazon's doing the improvements then. So if you're in a data center in a region for years, you can probably slowly see improvements and you hopefully see these like bigger drops like, hey, we've, you know, put solar on the roof, for lack of a better example, and now you can like see this dip in it. I guess this is more water, but you know, we've updated our water system, but like you'll see these dips. [41:00] Justin Brodley: Yeah. So you'll, instead of seeing, you know, a system that isn't a closed-loop evaporative cooler, and so it's using a lot of water, hundreds of thousands of gallons per day to, you know, a system that's now using 10,000 gallons a day 'cause it's now using evaporative closed-loop. That would be something you'd be able to see. So that's what I think you'll find in this. Now we talked about earlier about CloudWatch data moving between those different tiers and the one thing if you've ever been burned by this is that there typically was a 30-day minimum for transitions to S3 Standard IA or S3 One Zone IA and if you needed to move it back out of that, like you accidentally clicked the button, you need to undo it, you had a 30-day minimum or you paid a pretty hefty fee, which is always sort of annoying. Amazon has thankfully eliminated that 30-day minimum retention requirement for transitioning S3 objects to Standard infrastructure and one-zone infrastructure, allowing lifecycle rules to move data as soon as zero days after creation. This change directly benefits workloads where data cools quickly, such as backups, Log Analytics, and compliance archives, letting customers capture up to 40% storage savings without the previous waiting period. Previously, customers had to keep data in S3 Standard for 30 days before transitioning, even if the data was rarely accessed after creation. Implementation is straightforward through updated S3 lifecycle rules by console, CLI, or the SDK. [42:16] Justin: It's just a great quality of life. And I'm trying to figure out what caused them to make these changes, uh, because I've definitely inadvertently set things to these and then deleted them or moved them and then got hit with a fee. So what's causing them to make these changes now versus where it was? But hey, I'll take the win and move on in life. [42:37] Justin Brodley: I mean, I think they're, they're listening to customers, or the AI is saying, what are sharp edges that people complain about on Reddit or Twitter or things? And like, we could fix those things. And so yeah, allowing me to put something into infrequent access immediately for my backup is amazing. It's like, why store it in 30-day? Like, I have a Synology backup, I replicate to S3. The chances I ever need to use that within that 30 days, pretty darn low. So being able to move it right to infrequent access is a, a huge win. And, and today I do it with Glacier backup, so I don't even, I don't even have S3. But you know, that's something that changed recently as well, is I was, you're actually now able to write to Azure directly through S3. [43:16] Justin: You were able to write directly to S3 IA, but you, but if you put it in there, your lifecycle wasn't able to rotate down. 'Cause I definitely did that with, with stuff. So just a nuance there that, you know, but then I, at one point I set a lifecycle policy to delete after 7 days 'cause I forgot I started writing stuff to 1A, 1IA, and then I started getting hit with bills. I was like, that's not very nice of you. [43:41] Justin Brodley: No, definitely not. All right. Uh, Amazon CloudWatch is announcing Coding Agent Insights, giving engineering leaders visibility into AI coding tool usage and ROI. Integrating with Cloud Apps Gateway for AWS to pull telemetry from cloud code without extra instrumentation. Codex and GitHub Copilot are also supported. The feature is built on OpenTelemetry metrics and surfaces them alongside existing CloudWatch operational data. Letting teams correlate agent adoption with commit throughput, pull request velocity, and cost-to-output ratios by model. Practical use cases include setting proactive token billing alerts, tracking spend trends by department, and identifying which teams would benefit from expanded coding agent access. Availability spans all AWS commercial regions except the Middle East regions and Israel. Setup requires configuring Claude Apps Gateway to emit telemetry to CloudWatch per the setup guide. So I mean, this is nice if you've ever set up Claude to send data to an OpenTelemetry endpoint. You get some really interesting data. And so the fact that this is getting built in natively to the Cloud Apps Gateway, kind of surprised it wasn't there day one, but glad to see it's there now. [44:43] Justin: Yeah, I mean, I think it's just they know what to expose and now they're slowly exposing more and more things to people. [44:49] Justin Brodley: So, well, I think customers are demanding this too. They're saying we need more velocity, we need more telemetry into what's happening with cloud code, we need to know what people are doing. And so I think that's forcing some of their hand. [44:58] Justin: Yeah, I don't think they have an option. Because I think everyone's going right now, whoa, our AI bills are so high. What are we doing with it and how are we using it? And that's kind of what you're seeing now. [45:08] Justin Brodley: Yeah. I mean, the token maxing era, uh, was, was glorious for moments. Now it's over. So, yeah. CloudTrail now supports IAM identity-based filtering for network activity events tied to VPC endpoints, letting teams log only relevant traffic instead of every API call passing through a PrivateLink connection. A practical use case for this would be configure selectors to capture VPC access denied events only from identities outside a trusted allow list, which helps flag potential data exfiltration attempts while suppressing noise from known approved roles. This supports data perimeter strategies by combining user identity conditions with existing selector fields like event name or VPC endpoint ID, giving the security team granular control over what gets recorded. You can reduce both log volume and CloudTrail costs since routine traffic from trusted principals no longer needs to be logged, while still preserving visibility into anomalous or unauthorized access patterns. [45:58] Justin: I mean, it's great that you can actually start to select what you need. There was so much noise in there and so much stuff was like a needle in the haystack. Even once you followed all their guides and pumped it to Athena and then, you know, to your S3, you had Athena and we went down that whole path and then tried to search it, but still finding the deny log, you know, in there or finding the allowed of something that you were looking for that technically was allowed but you didn't mean to was annoying. [46:27] Justin Brodley: You know who's really sad about this? Who? Splunk. Because now, now you can filter the data. You don't actually care about getting to Splunk because that used to be, you know, these type of things were kills your SIEM costs. [46:39] Justin: Yeah. [46:40] Justin Brodley: Because of, you know, you're sending all this irrelevant data that you don't need to care about, but because there's no way to filter, you just ended up in a situation where you were logging everything. And, and there's, Sometimes it's helpful to have that data, but most time you don't need it. [46:51] Justin: That's how I feel sometimes about VPC flow logs. I'm like, it's great that it all is in there, but how often do you really need to be looking at every single— [47:00] Justin Brodley: Not very often. I mean, that's why I like Bump on the Wire, 'cause it was kind of like the, instead of turning it on for everything, I could just bump between the two things I actually cared about for whatever I was looking at. Mm-hmm. Which is great on the troubleshooting side, not so great on the security side. That's why they still want flow logs. But at least with flow logs, you know, you're not necessarily capturing everything. But then, you know, when you get into like the heavy duty packet mirroring setups where they're inspecting every pack it, I start questioning the value of those for sure. [47:24] Justin: Some security person said they needed it. [47:26] Justin Brodley: Yep, and they are paying through the nose to get it. So yeah, you're welcome. GuardDuty Investigation Agent is now in public preview, automating security finding correlation and investigation, cutting analysis time from hours to minutes by providing risk levels, confidence scores, MITRE ATT&CK mappings, and prioritized remediation steps. Investigations can be scoped to a single finding, a specific account, or an entire organization, and can be triggered by a console, CLI, API, or natural language prompt up to 2048 characters describing areas of concern. The agent integrates with the AWS MCP server, allowing teams to invoke investigations through natural language via tools like Claude or Kiro, and fits into existing pipelines, uh, like EventBridge to SIEM. So Lambda functions can enrich raw findings of the structured assessments before routing to incident response queues. This is distinct from AWS Security Event Incident Response, which pairs AI agents with human engineers for active incidents. The GuardDuty investigation agent is for on-demand assessments rather than incident coordination. Available at no additional charge during the public preview in the 10 regions it's available. [48:23] Justin: This sounds pretty cool. I'd be curious what it's gonna cost in the long term. [48:28] Justin Brodley: Yes. [48:28] Justin: Because it always worries me with that, but I think it can have a lot of value because, you know, there's so much noise, like we just said, there's so much noise out there. And how do you get the true, you know, what is this and how does it actually affect us versus, okay, there's a high critical, this thing is talking to this thing. Okay, but there's 16 layers between here and there. So it's fine. [48:49] Justin Brodley: Yeah. I mean, I like the idea of having this on demand though, cuz that's, you know, that's one of the problems with the other one is it is all the time. And so if you don't need it all the time and you're just trying to investigate an incident or you're trying to investigate something specific, being able to just interrogate or like as part of a potential PR review or, or something you're trying to do, you can now invoke this in that different model. And yeah, it's really a question of cost. Is this super expensive to use? Is this relatively low cost for the times you would use it?, or is it more cost prohibitive than just running the other thing all the time? Which I'm pretty sure is not the case. I think it was expensive. I remember Ryan going, oh my God, this is crazy. Amazon ECS now offers 3 bundled pricing tiers: Essentials, Pro, and Enterprise, replacing the previous model where deliverability features were purchased individually as add-ons. Each tier builds on the last. Essentials covers deliverability insights, Pro adds managed dedicated IPs, email validation, and inbox placement visibility. And Enterprise includes multi-region resilience, workload-level reputation isolation, and annual deliverability assessments. The bundling approach is aimed at simplifying procurement for customers who previously had to evaluate and purchase capabilities separately, with AWS noting the plans are discounted compared to à la carte pricing. Available in all ECS regions except the UAE and Bahrain. And this is interesting just in general because they've had the same pricing for ECS forever. [50:07] Justin: Uh-huh. [50:08] Justin Brodley: And then to have add-on pricing, it was getting pretty expensive to start looking at like, oh, deliverability insights. And then you turn that on, all of a sudden the price has changed. So I like the idea of bundling it as long as the bundling comes with those discounts that make up for it. I'm sure this also helps them get away from doing a lot of special private pricing for ECS, which is probably what they were doing a lot of because they didn't have this bundled capability. [50:29] Justin: Yeah, I mean, figuring out all the differences of ECS was kind of a pain. It was one of those things that It started off as a pretty basic service and just kept growing and growing and growing over the years. So it's great to see it. The big one I think that I personally saw and see that's going to bump people is people that want dedicated IP addresses. That's going to move people from the Essentials to the Pro, I think, pretty quickly. [50:53] Justin Brodley: Yeah. But again, if it's bundled properly, it's not a bad thing because you also want some of those other affiliate— if you have a dedicated IP address, you're going to want deliverability. Visibility. You're going to want those things because you're paying for the dedicated IP. You don't need to worry about warming the IP. You worry about blacklisting. You need to worry about all these things that you're paying extra for. So if the bundle is done properly, I think it's a, it's a natural, like, yes, I want dedicated IPs, I also need those things, right? We'll see what customer reaction is, but I think it's the right move. We'll just see. Yeah, but again, I don't necessarily know if you need, you know, multi-region resilience and workload-level reputation isolation. Like, That's, those are pretty, pretty in the weeds where, you know, inbox placement visibility and email validation dedicated IPs in the Pro. It's probably most people are gonna end up in Pro is my guess. [51:36] Justin: Yeah. I just always like how each region was its own thing. So if you ever got blocked in one region, you could just go launch in another region and Amazon would never figure that out. Yeah, definitely never had to do that to bypass, you know, reaching out to the ECS team, get them to unblock us because we were doing what we were told to, we were allowed to do as part of our customer contract. Definitely never had an argument with them. [51:57] Justin Brodley: Yeah, never, you know, never had an argument with them about bounce rates before. [52:00] Justin: Yeah, had to, you know, reach out to the ECS PM in order to get them on a call to apologize for blocking us. Definitely never happened. [52:09] Justin Brodley: Never happened. All right, moving on to Google Compute Cloud. Notebook LLM has been rebranded. Could you make a guess of what it's been branded to, Matt, knowing how Copilot takes over the— no, no, that's Microsoft. This is Google, so it's going to be Gemini. Gemini Notebook is now what used to be known as Notebook LLM. Gemini Notebook is reflecting its expansion beyond a standalone research tool into deeper integration with Gemini app and Google Search, following adoption by over 30 million users and 600,000 organizations since its 2023 launch as Project Tailwind. I mean, basically, Gemini Notebooks is the closest thing they have to Cowork. And since, you know, ChatGPT now has a Cowork version, and Anthropic, of course, had Cowork forever, they invented it. [52:50] Justin: It. [52:51] Justin Brodley: Um, this just made sense to me. This is, this is a no-brainer. Like, yeah, so you take Notebook, you turn that into more of a co-working type solution on top of Gemini Enterprise, you rebrand it, and now everyone thinks it's all one product. Which Google desperately needs brand marketing help on this stuff. [53:07] Justin: I mean, I actually really like Notebook LLM. I use it a lot for like personal things where I'm just like, I only want you to understand this set of data. I kind of dislike that they're been moving it more and more into the Gemini ecosystem because it was a good way to just have it focus on just this versus inadvertently pull all this other data. [53:27] Justin Brodley: Yeah. Yeah. I mean, I like it. It's also the thing that competes with us as a Cloud Pod, you know, because it creates podcasts of anything you give to it, which is kind of neat, but, uh, not as good as us. I don't think, Matt. I don't think we've been replaced yet. [53:38] Justin: We get our satire and our cynical, cynicalism. [53:42] Justin Brodley: Yeah. I mean, Gemini's not gonna tell you that Oracle's terrible. It's not gonna do that. Why would it do that? [53:48] Justin: It might. [53:49] Justin Brodley: It might. [53:50] Justin: Probably not. Bolt started getting sassy with us the other day. So, you know, AI is slowly learning. Yeah. [53:56] Justin Brodley: I mean, it's, it's picking up on, uh, you know, things in our show notes. And so it knows, uh, Ryan is a big fan of the Eagles, for example, and it'll tell you that all day long. [54:04] Justin: And that we all really love Kubernetes and thinks everyone should use it everywhere. [54:08] Justin Brodley: Yes, exactly. Cloud Run now supports automated failover for multi-region services with two new capabilities. First, a readiness probe for instance-level health checks and a service health aggregation across regions exposed via serverless nginx. When paired with Global External Application Load Balancers, traffic automatically shifts away from unhealthy regions within seconds, removing the need for manual incident response during regional outages. Two deployment paths are supported: Global External Application Load Balancer for public-facing apps and a cross-regional internal Application Load Balancer for private VPC traffic. This works best in an active-active configuration with read-write heavy workloads that synchronize data across regions. Teams still need to handle redundancy at the database layer separately using services like Spanner, Firestore, Cloud SQL, or Cloud Storage in a multi-region configuration. This is available now in all Cloud Run regions at no additional feature cost, and customers only pay for standard CPU and memory charges for running the readiness probe. [55:02] Justin: I like multi-region, but I also feel like and this is where my knowledge of and lack of really running production workloads in GCP has— the whole concept of they're less region-specific, I thought. So I guess I'm just a little bit more confused about this announcement overall. [55:20] Justin Brodley: So there are things that are very globalized inside of GCP, things like the network and IAM policies and things like that. But the actual compute stuff that actually runs on computers is very regionally specific. So when you configure Cloud Run tasks, you configure them in a specific region to run there, and then you have to handle either run those tasks in multiple regions and you route traffic to them, or now in this capability you have the failover option, which we did not have before. But yes, anything once you get beyond the basic global services, it is very regional. [55:53] Justin: Got it. [55:55] Justin Brodley: Google's giving us 3 new Flash models, the 3.6 Flash, the 3.5 Flashlight, and a specialized 3.5 Flash Cyber for security use cases. All targeting improved efficiency for production AI agents. 3.6 Flash uses 17% fewer output tokens than 3.5 Flash while improving coding, knowledge work, and multimodal benchmarks like OSWorld Verified, an MLE benchmark. Pricing is notably lower than 3.6 Flash at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Still not as cheap as KIMI, reducing overall cost for agentic tasks compared to 3.5 Flash. 3.5 Flash Lite pricing is $0.30 per 1 million input tokens and $2.50 per 1 million output tokens running at 350 output tokens second, making it suited for high-throughput workloads like document processing and agentic searching. Uh, the 3.5 Flash Cyber is a specialized model fine-tuned for vulnerability detection and patching deployed through Google's CodeMender agent, which we'll talk about in a minute, using multiple coordinated model instances to generate consolidated security reports. Access is restricted to governments and trusted partners by a limited pilot program due to the dual-use risk of cybersecurity-focused AI. They just didn't want to end up like both Mythos and ChatGPT. GPT. [57:01] Justin: So yeah, they were just bypassing that, or they talked to somebody in advance, one of the two. [57:05] Justin Brodley: Exactly. So yeah, I mean, if you're into the Flash models, uh, which is really— it kind of undervalues what they do because the Flash models are pretty darn good. I mean, most of the things I do with Gemini are on Flash. You know, you don't need to use the Gemini Pro models very often, although like image generation, I use typically the more Pro image models, uh, because they listen to me better than the Flash ones do. But So nice to see these are getting updated once again. [57:31] Justin: Hey, these improvements are always good, but I'm really curious about KIMI. I feel like it's the big winner for me. [57:37] Justin Brodley: Yep. If you haven't tried out KIMI, you definitely should check it out. It's pretty great. All right. CodeMender, born from Google DeepMind research, moves into preview as a managed AI agent that scans, verifies, and remediates code vulnerabilities available via Gemini Enterprise Agent Platform. Helm, or as part of the AI threat defense capability. The agent goes beyond static analysis by building and running proof-of-concept exploits in a customer-managed sandbox to confirm vulnerability is actually exploitable, reducing false positives and alert fatigue before generating a fix. Remediation is delivered as code dev for developer review with an LLM as judge, step checking that patches don't break existing functionality, and developers retain approval control before anything is committed to the repo. Support common— this is a support common languages including C and C++. C++, Go, Java, Python, Ruby, Rust, and TypeScript, and integrates into CI/CD pipelines, VS Code, Antigravity, or a CLI client for local development workflows. This follows a multimodal approach, allowing teams pick models based on cost, speed, or scanning depth, with third-party Frontier model support planned later this year. Specialized version with Gemini 3.5 Flash Cyber is limited to select government and trusted partners access initially. Integrate with the AI threat defenses using Wiz to orchestrate the workflow, calling CodeMender to scan code and enrich findings with the Wiz security and trigger Wiz Red Agent for AI pen testing to prioritize the highest risk issues. Early customer quotes came from Salesforce, Robinhood, and Palo Alto. If you're curious in the article, check those out as they are very positive, of course, because they're being quoted by Google. [59:02] Justin: I mean, I think that you're starting to see Wiz get more and more now embedded in the entire Google ecosystem, which is nice to see. I mean, I think it's going to be really good for security overall. Letting Wiz see a lot more and more things. So I understand that's probably the least important part of this announcement, but it's the piece that I really think about. [59:24] Justin Brodley: All right, let's move on to your best friend, Azure. Uh, first up, they're, uh, expanding Azure infrastructure with 3 new AMD-powered VM families: the HD v2 for data processing, the HX v2 for electronic design automation, and the MD MI455X v7 for AI inference, all built on AMD Kubernetes, Helios platform. I mean, cool, I have new big servers. So if you're into those things, uh, new instances for you to not remember the names of, because I never remember the names of these things. [59:53] Justin: I don't remember these ones. I mean, I get their naming convention makes sense if you understand and you have the trans— you know, the translator for it, but I don't ever want to use these things. I give them enough money, I don't need to give them more money. [60:06] Justin Brodley: You don't want 500 physical 6th gen EPYC cores? With 4 terabytes of RAM and 32 terabytes of local NVMe storage and 400 gigabytes of networking. I mean, come on, that's a massive box. [60:17] Justin: I'll run, uh, my Home Assistant on it. [60:20] Justin Brodley: Think how well SQL Server can run on that. You could, you could back your Home Assistant with SQL Server. That's how it works. [60:25] Justin: Just don't think about, uh, the pricing of the Microsoft SQL licensing for it. [60:30] Justin Brodley: Yeah, don't just use the Developer Edition. It's fine. What's the problem? Uh, and just for a reminder, uh, Skype Business, uh, 2015 and 2019 Extended Security Update program is not going to be extended beyond October 2026, closing out the period 2 of the ESU that followed an earlier one-time extension. Organizations still running Skype for Business 2015 or 2019 in production will receive— [60:55] Justin: You really don't need to keep going. [60:56] Justin Brodley: I know, I will save no further security updates, but I'm just saying you also don't need to reach out to me about any future job opportunities because if you're still using Skype for Business 2015 in 2026 or 2019, I'm sort of wondering what you're doing. [61:10] Justin: I mean, I just saw this come through and I was like, I just thought this was funny that it's still— [61:13] Justin Brodley: I thought it died like 3 times already. I, I feel like every 6 months we talk about Skype for Business dying. I mean, I mean, it, it's not actually dead either, which is the worst part of this. It's part of Teams. Like, yes, the official product SKU is dead, but like to move off of this SKU is just a matter of moving to Teams and it's done. It's sorted. Like, what are, what is holding you back? Back that you can't get off of Skype for Business 2015 and move to Teams. [61:39] Justin: And it's even in— like, you— if you ever see like the cards say like something dot Skype, and there's like Skype is so far embedded in there. And the couple massive companies I've worked with that were running Skype 5 years ago, I know in the last couple years it's moved off. So like, who's left? [61:59] Justin Brodley: Someone who's really picky. It's probably the government. It's always the government. [62:02] Justin: Yeah, it's probably some like, you, out of band, you know, secure network or whatever. [62:08] Justin Brodley: Dark zone with, you know, spooks who don't wanna move. I don't know. Well, anyways, if you're still using Skype for Business and you're listening to our podcast, uh, get off of it. It's over. Yeah, it's done. Move on. And finally, our last story of the week, uh, from our friends at DigitalOcean. Uh, they're apparently raising on-demand prices for NVIDIA and AMD GPU droplets effective August 1st, citing strong demand for GPU capacity. Existing customers must destroy droplets before that date if they want to avoid the new rates. Billing changes will apply to any active workloads running on or after August 1st, with updates charges appearing on the September 1st, 2026 bill, giving customers roughly a month lag before seeing the full impact. Uh, reserved 12-month pricing is also increasing, but customers currently under contract keep their locked-in rate. So, uh, quick, go lock in a rate, which incentivizes existing customers to consider extending contracts before that change goes into effect. And this is just a, you know, a reality of the continuing pressure on the compute market, uh, for memory and GPU scarcity. Uh, you're going to see prices going up across the board on GPU and memory-dependent services like this. So unfortunately, it's happening to DigitalOcean, but it's also happening to everybody else too. So they're not the only one in this camp. [63:17] Justin: Yeah, I think you're gonna see it pretty much everywhere. And you can even see the pricing of every new processor is going up as they do it. It's no longer like, hey everyone, please move from the M1s to the M3s and we give you a 5% discount every single time. Now it's like it's either flatline and they're like, well, it's flatlined, but at the same point, you know, you're getting an extra 5% out of it. So the, the story has changed over the years and here you're just getting the increase cuz it's just more expensive. There's only so much they can do about that. [63:47] Justin Brodley: Yeah. Well, look now, you have a week before August 1st, so get to it. All right, Matt, it's been another fantastic week with you here. We'll see you next week here at the pod. [64:00] Justin: See ya. Another week of cloud news wrapped up. Colt will collect the news. Justin will get the notes. Jonathan will write some code. Ryan will watch the perimeter. [64:12] Justin Brodley: And Matt will review. [64:13] Justin: Reluctantly watch Azure. Till next week for AI, Amazon, Google Cloud, and Azure. And hey, maybe even Oracle, who knows? Check out thecloudpod.net for our newsletter. Join our Slack, message us on socials, or leave a review.