# 371: MrBeast Bets on Gemini for Survival Duration: 73 minutes Speakers: Justin, Matt Kohn Date: 2026-09-17 ## Transcript [00:07] Justin: Welcome to The Cloud Pod, where the forecast is always cloudy. We talk weekly about all things AWS, GCP, and Azure. We are your hosts, Justin, Jonathan, Ryan, and Matt. [00:18] Matt Kohn: Episode 371, recorded for September 8th, 2026. MrBeast bets on Gemini for survival. Good evening, Ryan. [00:28] Justin: Good evening. [00:29] Matt Kohn: Only took me 3 tries to do that for the first time. [00:31] Justin: Yeah. Well, they edit that out if you don't say anything correctly. [00:34] Matt Kohn: You know, you gotta make fun of yourself, otherwise who's gonna make fun of ya? Yeah. Shall we jump in since both of us, uh, have sick things around our house? [00:44] Justin: Yeah. [00:45] Matt Kohn: Yeah. [00:45] Justin: So it's just the 2 of us this week. We've been left on our own, suckers. So we're going to our, uh, usual alternate format of round robinning the news and Being hilarious. No, not, not that part. [00:59] Matt Kohn: Definitely going on digressions. Anyway, we have, we get to skip follow-up general news this week and go right into AI is how ML makes money. Announcing the Databricks Big Book of Agent Ops. Databricks released a Big Book of Agent Ops, an ebook framework covering people, process, tools needed to move AI agents from pilot to production. Positioning agent ops as the operational layer between the existing MLOps and LLM ops practices. Because everyone needs more ops in life. The guide outlines 6 chapters spanning architecture, 7-phase roadmaps, evaluations, and feedback loops, DevOps-derived practices for non-deterministic systems, planning frameworks, stakeholders, and RACI governance, models because nothing can be real without a racing chart, guys. Customer site includes FactSite, text-to-code agents receive 44% accuracy improvements after moving to a full agent system. DXC reduced platform total cost of ownership by 30% after migrating to Databricks. The framework centers on 3 existing Databricks platform components, MLflow for evaluation and tracing, Unity Gateway for model and tool traffic, and Unity Catalog for governed data and access control, positioning these as technical foundations for scaling agents governance across organizations rather than managing controls for each individual application. I feel like nothing they're saying here is going to be life-changing, but I think what it is is really just laying it out in a way that people can follow it. Because these things are, at least from what I've read about it, is the general of what we all, what we talk about, what, you know, all the things are talking about out there, you know, controlling it, training, et cetera, et cetera. But I think it's just laying it out in a, in a nice easy way for people to actually read it. [03:00] Justin: Yeah. There's so much, you know, confusion and, you know, gray areas in the, in the market. And, and when you read through documentation, it's really easy to get lost on like, oh, what are we talking about? An agent. That is like my coding agent, or are we talking about an agent that's part of an application that's in runtime or that, you know, and it's just, you, you apply these different models of operational, you know, guidance in so many different levels. And so like, I, I do kind of like the sort of real world example of this. Uh, I don't know if I agree with their position of, you know, sort of like the MLOps, and this is like another sort of parallel in there, but I, I do think that it's kind of, good to have the sort of agent architecture laid out just because I think that people don't think about these things as, as an application exactly. And you know, this is sort of documenting the whole cycle and of, you know, development, if you will, quote unquote, which I think is neat. And you know, I think what we sort of will have, we're sort of having to invent this as we go, cuz it's, we're not treating it like application code yet. But we're relying on it like our production applications and rolling it out. But it's one of those things where I think we need a lot more thought around, we need a lot more established patterns and paved roads to do it correctly. [04:20] Matt Kohn: I don't think we have, I think I agree with you wholeheartedly on we're treating these things not like applications, but they are. But I also think that because we're not treating them like them, we're trying to now build this whole other framework around it But even, you know, the article, our show notes and everything kind of all allude to the fact that, you know, we're trying to follow DevOps derivative practices, which to me is just following software development, you know, and your SDLC process and everything else. And I don't think you're really, you know, I haven't read the book. I have it as my homework. I have a few long flights in the next couple of weeks. I'm going to read, you know, gathering material to read on, you know, and this is going to be one of them because I don't feel that. We need to redo our entire paradigm of how we do things. I think there might be small tweaks, but we should just follow an SDLC. Yeah. [05:13] Justin: I mean, it's going through and it's just really convoluted. Like I was going through, you know, the ICE, um, SQL to text application that they mentioned in here just because, you know, it's my old alma mater. I don't know if you can say that about ex-jobs. Anyway, uh, and you know, like it's, they go through and they talk about it like it's this big thing, but it's really like you, when you read through it, it's like they're, they're kind of just, it's not really, they haven't really put together a cohesive sort of system that would do, that would fit in an SLDC model, right? They're sort of using different parts of Databricks and it's a, you know, it's a big publicity blog, so that's of course gonna be in there. And they're largely just sort of like, you know, interacting with different elements of the Databricks platform, right? Like it's not smooth, it's not polished. And it's kind of like, I was trying to figure out like, depending on how they're using this, right? Like it doesn't really go into like sort of the business outcomes that Intercontinental Exchange is looking for, you know, with this, with their text-to-SQL app. But it's, it's obvious, you know, like I don't want to write SQL ever again. So I do understand that, you know, the, the overall text-to-SQL makes sense, but it's also sort of, you know, like what are they going to do with the output of these queries? Isn't really mentioned. Even though there's a section specifically about impact, but they're just talking about the accuracy and execution matches, which I thought was funny. But you know, it's one of those, like, I, you know, we need to have more comprehensive ideas around these things and it's, if this is something that they're, you know, selling or if this is anything more than like an internal sort of BI sort of tool, it needs to be fleshed out. And I think what, A lot of places in the industry were not really rolling it out with that level of sophistication. [07:00] Matt Kohn: No, most AI is just rolled out as fast as they humanly can to make the shareholders and everyone happy being saying, you know, we have AI and, you know, a great example was I read somewhere that, you know, a product team said, oh, we need to have, uh, or like, I think it was like a CTO, we need to have AI in our search and they did it. And the product team was like, why? Our search works. This is an area we need to focus on. You know, and then the article came out later on where saying that customers actually hated the new search 'cause it was less accurate and everything else. So AI, you know, in general needs to be put in the right place in the right time, not just thrown at everything. And that's where I kind of feel like a lot of the market has been is throwing it and see what sticks. but now we're starting to get to the point when you can't just do that as much. [07:52] Justin: All right, moving on. OpenAI announces the rollout of GPT-6 Astra models. They're rolling this out in phases, starting with companies in its Daybreak cybersecurity program before a wider availability on ChatGPT Plus, Pro, and Business plans. The OpenAI API in AWS is, is coming in a few days. Azure is the first OpenAI model to hit the company's internal critical cybersecurity threshold, prompting additional safeguards and restricted access to its most advanced capabilities, which, you know, is good considering this is following a pause in research and training after the previous two incidents where the models breached containment and accessed Hugging Face's systems. [08:35] Matt Kohn: I still have questions. Yeah. Like, what was the real problem here? Was it the marketing or was it the fact that you don't know how to properly say a box? [08:44] Justin: Something. I mean, they're— [08:47] Matt Kohn: the— [08:47] Justin: how it got out is pretty sophisticated. [08:48] Matt Kohn: I don't— yeah, the whole art factory thing. Yeah, yeah. [08:54] Justin: Uh, OpenAI states that Astra shows improvements in computer use, software engineering, multi-step workflows, task boundary adherence, and understanding user intent compared to the previous models, positioning it for enterprise delegation of more complex tasks. It shows memorable gains in benchmarks, scoring 72.6% in OSWorld 2.0. In about 40 minutes per task versus 65.7% in 75 minutes for GPT-5.6 SOL. That's roughly 47% less time per task, which, you know, great, and probably 4 times the token burn. But OpenAI defines their critical threshold in cybersecurity bench framework as it scored 100% on Exploit Bench and 42.4% on Exploit Gym compared to 78.5% and 30.3% for GPT-5. 5.6. Astra also discovered two previously unknown zero-day vulnerabilities, which OpenAI is disclosing to its maintainers. And this capability increase has prompted additional safeguards. Good. Restricting on generating proof-of-concept exploits. Alignment testing shows Astra deviated from authorized target 0% of adversarial cases versus 48% for GPT-5.6 without the production safeguards. It never attempted to bypass the Codex auto-review denial, even when it was made deliberately available. Azure is also reported to be 3 times less likely to misrepresent its own capabilities. Thank you. And it introduces new context preservation method, allowing the model to retain search notes across long sessions instead of relying solely on compaction summaries, useful for extended debugging or large refactors. It's available now as an ex— experimental config option and becomes the Azure default in coming weeks for enterprises. Zero data retention is supported for eligible API customers, and admins must manually enable Astra since it's off by default at launch time. [10:51] Matt Kohn: There's a lot of interesting things that are direct attacks, I feel like, at Anthropic in this, you know, the off by default, the, hey, we're not collecting data because Doesn't still Fable send back all logs for, and they retain them for like 30 days. You know, so there's a lot of like little things in here that I think are big. The new, I'm, I'm really curious about the new compaction process that they have to reduce, you know, your token usage, because that's kind of big, you know, definitely have had long conversations, especially when writing like statements of work or long documents where I'm going back and forth with it for. Days, you know, whenever I get 15 minutes here, 15 minutes there, and I just see every time I restart, 'cause the cache misses and everything else like that hit really hard on a long conversation. So I said it in the Slack channel the other day, it's definitely gonna be, you know, on my list of things to do is really go play with Codex and start playing with the GPT models again, 'cause I feel like they've really started to gain some traction. Yeah. [11:57] Justin: Anything with changing the context, compaction is good. That's, that continues to be the bane of my existence. And I'm, I'm, you know, where that's where I'm tweaking all of my work, internal workflows for, for my day job, like whenever I'm completing work and I did a full analysis of where my token spend was in order to sort of see where I could get efficiencies and all of it is, you know, a lot of it is just those long, running conversations and the amount of context in those windows. And so that context I really want, um, because it stops you from having to repeat yourself. And, you know, and it's a little bit easier to have in that context than it's like, oh, go read this Markdown file. Or, you know, and having to— and knowing when you need to actually get something, you know, output it to memory or that kind of thing, it's a little complicated. And so like, it's— I think that this is going to be an area that we see a lot of improvement in because it is sort of still very painful, I think. And nothing makes me more infuriated than, you know, context being compacted and this thing we've been working on gets lost in the summarization and then it gets, you know, whatever gets reintroduced into the thing. [13:03] Matt Kohn: Like, no, it's the long conversations I think are the real killer. You know, we're gonna digress a little bit down to like token usage. You know, you and I are probably the same. We're in and out of meetings all day. We're trying to do, you know, I mean, quote unquote, so real work on the side to our day jobs. And, you know, we're trying to get stuff done and, you know, you lose track, you lose 20, 30 minutes, you come back, et cetera, et cetera. And all of a sudden, you know, I see like my tokens jump up and go from there and that's never good. You know, so trying to kind of get those pieces where you get the cash hits a little bit better along with the less, you know, trying to get back your history and compaction and everything. Those are gonna be big wins. It's interesting how much they're putting into it. So I'm curious to see if in 3, 6 months we keep seeing improvements, 'cause this is also where tokens gets used and tokens are how these companies make money. So it's kind of a double-edged incentive here of how much do they want to improve? But at the same point, if their competitors improve, they have to. So somewhere in there's a win, you know, I think they're seeking better performance. [14:11] Justin: Like you see the context window token size, you know, increasing with the models not really going down. So I think it's more, I don't think we've hit sort of any kind of point where they're gonna try to balance it out to increase their margin or something like that. But you know, it is, I think, a big impact and it has a, you know, it impacts how people design their interactions, especially for like coding workflows or, you know, these people that are doing like Hermes things in their automation, you know, in their homes, like you need a lot of gates and checkpoints where you're outputting sort of state because the, you can't rely on, you know, the, the context being accurate a lot. So you, you know, a lot of loops are generating just markdown files and text files. It just blow it up. Apparently GitHub is taking it all down because it's all just markdown. [15:01] Matt Kohn: I mean, my home lab downstairs has a lot of markdown files associated with this point. 'Cause at this point I just say, hey, go, go help me with fix this oddity. Or, you know, I have some sensors in my freezer and whatnot. And it's amazing in the summer, you know, up in New Jersey, we'll see, you know, the humidity is up, but I never, I've had it up for a year now and all of a sudden I see these big spikes on the humidity. So I've been like playing with like how far back do I show the charts when the humidity gets too high? 'Cause Kubernetes, humidity could be a sign that your freezer door's left open just a crack. So you start to get into these weird things. So I'm sure I have a lot of Markdown files there around these things. [15:41] Justin: We've graduated from the days of piles of YAML to heaps of Markdown. [15:46] Matt Kohn: Yep. Yay. You know, I think I like Markdown better, but I'm sure there's a battle somewhere online about Markdown versus YAML. I still say it all beats JSON, but I hate many things about everything. [15:58] Justin: I think they're different functions, so I think we're safe there. But I actually prefer JSON over YAML, so that's my hot take on that. [16:06] Matt Kohn: So I don't know, I've been thinking about that a lot over the years, especially with CloudFormation change and everyone was like, there's YAML. And I was like, but JSON worked. But I also have fought JSON for hours finding the one curly bracket that I want to bang my head against the wall for missing for 20 minutes. [16:22] Justin: So, you know, yeah, but now there's like inline parsers that tell me where I messed up. [16:26] Matt Kohn: So yeah, that was a lot better. Yeah. Or auto-formatting even with a broken one. So you can see the tab. Yeah. All right. On to security. As we've digressed into how JSON files are properly formatted. BGP hijack, hijack infecting networks caused by comedy of errors. That's not funny at all. We should have figured out a show title based on that somewhere. [16:49] Justin: Yeah. Anyway, I thought about it. [16:53] Matt Kohn: Attackers use BGP hijacking. Against Hetzer Online to seize control of soft malicious IP spaces, then push malicious updates disguised as legitimate software for Visualizer, a visualizer management platform used by hosting providers and data center companies. Two separate failures enabled the attack: weak routing security configuration configurations at Hetzer. I don't know how to say that properly, so someone can yell at me later on— that allowed for IP hijacking and soft malicious— no, no, not using code signing, really bad in life— to validate update packages, meaning malicious updates would have been rejected by the client. Hijack ran intermittently over a 33-hour window. Hydra reclaimed the address space after 12, but the attacker repeatedly hijacked it and took nearly 10 hours to respond the second time, extending the exposure window. Lambda. They will not confirm which servers were compromised and advertising customer— and advising customers to treat all installations potentially affected, illustrating the difficulty of scoping damage after a routing-level supply chain. The incident highlights two longstanding infrastructure weaknesses worth discussing: the industry's slow adoption of RPKI and other BGP security measures. And the risk of software update mechanisms that lack crypto— crypto— cryptography verification, both of which are basic and well-known mitigation methods. I feel like the— we'll leave the BGP because BGP hijacking and mistakes that have been issues for years. There's fun Wikipedia pages if you ever want to dive down it. For example, when like all traffic routed through China for like 10 minutes on the internet, or when like, I think it was like Pakistan accidentally took down the internet by publishing 0.0.0.0 and then it routed everywhere. Or when Facebook locked down, locked themselves out of their own, I don't remember that, like they accidentally pushed a BGP update that zeroed out their entire BGP thing and like they couldn't get into their data center physically, but we had, was it last year? Must have been a year and a half ago. Notepad++ that had kind of the same type of thing with the software updates where somebody got in and potentially pushed out an update and not having your code signed was a big piece of it. And it feels like at this point, at least you are far more of a security person than I am, but I like to think of myself as a fairly secure person. Your code should be signed and it's not that difficult. You know, there are code signing ways, whether you pay for them. I know DigiSign has it, but also like Azure has one built in. I don't know. I assume AWS does, but I've never looked at it. Like there are online, fairly cheap code signing, like ways to do it. And it's not too difficult to sign it. Maybe you don't need to sign every package, depends on how it works for you, but like you gotta sign these things at this point. Yeah. [19:58] Justin: I mean, it's, you know, supply chain attacks suck, right? They're, they're really difficult. And I kind of, I feel both ways about signing just because like if you, you know, there's also a barrier of entry for, you know, getting your stuff signed by the, the Apple developer, uh, sort of App Store, you know, signing that they do, right, to, to validate these things. And so it's, it's a little bit, you know, gatekeeping in that sense. But, you know, it depends on your relationship too. Like I'm not familiar with this virtualizer software or Or softculeshous, softculeshous, like, I don't know what the model is for, you know, for the downloading the software and that kind of thing. But it is sort of whenever you, you know, you're presented with something that looks legit and you download, you know, some application, run it and run it on your system. Like that's, it's a huge attack vector cuz you're, you know, it's like, oh, it's asking me for my password on this thing I just told it to do. Then you're giving it full access to your system. Not very good. So it's, you know, like, you know, I feel like it's, I do think code signing is a good idea, but this is more like, I don't know if this is an application or, or if this is more, you know, mutual TLS space, or I'm not really sure what the thing is, but I think the bigger one is actually the, the BGP hijacking because there's the validation and the signed the signing of the advertisement within BGP has been around for ages, but no one's adopted it. You know, BGP is just this magic that makes the internet work. And so people aren't, you know, it's just kind of at this foundational layer and people aren't paying attention at it. So they, you know, you can't really rely on any kind of distinction if someone does hack your, um, you know, it's sort of like circumventing the advertising. Which is what happened in this case. And so it's like, well, as more and more of this happens, I imagine there'll be adoption, but it's, it sucks that it's sort of has to be through this outage and incident that's adopting the sort of validation. [22:06] Matt Kohn: Yeah. I mean, BGP is one of those core things on the internet that I'm not even gonna say the average person, I may go with even decently good software developers probably have no idea about BGP. They know DNS, they know there's like— BGP is a networking thing inside and out, unless if you're doing even on the cloud Direct Connect. [22:29] Justin: I mean, you're not running BGP inside your own network, MLS apparently, or Azure with their latest announcements. [22:35] Matt Kohn: Yeah, well, we'll bypass those. Yeah, like most people aren't using BGP. And you know, what's the other one I've used? Uh, OSPF. What, uh, it was another routing one that I've done between like my firewall and two switches. But like BGP isn't something that people use, use. And it's so like, it's designed because the internet's designed to be open and, you know, trustworthy, but it's not a trustworthy place anymore. And so you're now retrofitting all these old systems with security, which you know probably better than me, right? It's really always fun to try to retrofit security into a running system, especially at the scale that the internet's on. You know, so BGP is just one of those things that it's hard to do and making changes requires, you know, massive soft, you know, hardware providers to take the updates and make it and put it out and then rollouts at, you know, core internet providers on these core internet BGP devices. And it's, it's a long tail process. I know for years they've talked about securing BGP and I know they've done some, but. Clearly there's still more to do. And I don't know much about the RPKI, but I bet you it's going to get, like you said, more visibility quickly. [23:53] Justin: Yeah, I mean, it's anything where you're having to coordinate at the internet scale, right? Because it's— I, I don't know if BGP is all that complicated, but I really only understand it at a high level myself. So like, I'm not a network engineer, so it's sort of like I get the basics of it, which I think is more than your typical developer, but you know, like it's, maybe it is much more complicated than I think, but it's a lot of different, it's public participation, right? It's public coordination of every endpoint on the internet. And you know, it's how routing all the different global endpoints, you know, is accomplished is through those BGP advertisements and the coordination there. Like, and so like if feed gets hijacked or, and I wonder if if the article has more, or I wonder if we can find more specific details on how the hijack took place, but, 'cause I know how you do it to yourself, like Facebook, like I know how they did. [24:46] Matt Kohn: Yeah. Yeah. [24:47] Justin: But I don't know how, I mean, I guess you just start advertising a smaller subnet, but you just, you advertise or smaller. Yeah. Not subnet, but the smaller sort of neighbor. But yeah, I don't know. Like that's, that's difficult. [25:00] Matt Kohn: I posted in the, uh, in the Cloud Pod general Slack channel. You know, it's one of those diagrams of the internet and how everything's balanced on each other. And, you know, the bottom says DNS. And as we're talking about it, I just think more and more of, it's not really DNS. It's really, you know, below that it's BGP. Yeah. It's one side of it, not just DNS, but. [25:22] Justin: There's a whole underlayer that we don't even discuss, you know? [25:24] Matt Kohn: Right. [25:25] Justin: That's one of them. Like all the ICANN, like IP stuff, like how all that works. Like it's just this internet cabal. Like it's so crazy. Yeah. [25:34] Matt Kohn: Shall we move on and actually talk about the clouds? [25:36] Justin: Let's move on to AWS. AWS Lambda now supports SnapStart for container image functions. SnapStart extends to container images, cutting cold start times from several seconds down to sub-second by caching a snapshot of the initialized execution environment and resuming it on invocation rather than initializing it from scratch. This closes the gap for customers who package their functions as container images to meet organizational container standards or to bundle larger dependencies. Previously, SnapStart was limited to manage runtimes like Python, .NET, and Java in zip deployments, but apparently it's expanding out to more broader dependency bases. It's useful for latency-sensitive workloads such as ML inference and interactive APIs where startup delay directly affects the user experience or response time SLAs. SnapStart is available in all commercial AWS regions except Asia Pacific, New Zealand, Asia, Taipei, and it can be enabled via the Lambda API console, CLI, CloudFormation, SAM SDK, or CDK for new or existing functions. For AWS-based images with Java, Python, or .NET, the experience matches the existing SnapStart behavior for .NET. For .zip archives and other types like Node.js, Ruby, or custom base images. Although those custom base images require additional configurations per the developer guide. Pricing details and usage base are listed in the AWS Lambda pricing page under SnapStart pricing. [27:10] Matt Kohn: It's a great feature for them to add to containers. You know, I definitely get why they did it and they probably did a lot of testing with the— I'm going to call it default runtimes of Python, etc. And I'm sure they kind of just cleaned up the back and this was just, okay, now we're ready because under the hood it's all containers. So I think they just had to kind of get all the caching and all the different pieces of the scale with what they knew. And then from there they were able to kind of expand onto all these custom containers. So I think it's a great quality of life improvement, especially if you need that. I still question if you're running 10GB containers on Lambda, why you need to boot it up in sub-milliseconds, but you know, you do you at that point. Yeah. [27:53] Justin: Yeah, that's, that is a pretty big function. It is kind of, you know, it's interesting. Like I haven't really had a latency-sensitive workload that I've ran in Lambda. I've done more like batch processing and sort of state machine sort of functionality that didn't have direct user impacts, but I've definitely been, involved in other workloads, like sort of ancillary. And it is sort of this, you know, I don't know, it's like a catch-22 a little bit in terms of like, you want the serverless functionality, but having to do all the initialization of your application and dealing with cold start is problematic. But you know, you get true scale zero. So it's, you sort of have to balance between the two. But I do like this, um, you know, fixing that, I think it's a good option for people. Um, I wonder how they do it. Like when I think about the scale of how, how they must, it must just be caching layers. You're per workload. Like they're just having a, like, I mean, it's got its own pricing model. So I assume they are just storing this alongside your runtime somehow. Yeah. [28:58] Matt Kohn: I assume there's some like extra caching that it does near for X number of time. Maybe it. I'm oversimplifying, but I'm thinking almost like, you know, Redis running next to it that caches the image on the way in, has it, and then the next one's just starting to load, you know, just like CloudFront, you know, like, oh, it's just a, you know, at the runtime. [29:18] Justin: That's crazy. [29:19] Matt Kohn: Yeah. [29:19] Justin: I mean, just like, it's one thing that, you know, like a key value cache in Redis, but this is a full execution environment with all the details, all the dependencies loaded into memory. So it's like, it's a big memory chunk, I think. [29:33] Matt Kohn: For other things like what's the pricing on? Is it priced on storage? I was trying to load it. Cache per gigabyte second and per restore for every gigabyte restored. So essentially it's going to be if you have a 10 gigabyte one, it's— I mean, it's not a lot, but if you're restoring it, one 10 gigabyte restore is still sub $0.01. Below, still zero cents, $0.0013. [30:03] Justin: Yeah, I mean, and, but it's then, you know, it depends on your scale, right? [30:06] Matt Kohn: If you're restoring it thousands of times, you know, a month, it adds up. [30:10] Justin: Well, and you think if you've got that set for your workload, it's every invocation is a restore, so yeah. [30:17] Matt Kohn: Well, bringing things from the future to now, Amazon Linux 2027, yes, 2027, that took me a few times myself. Is available in public preview. Amazon Linux 27 enters public preview built on AL23 baseline with kernel 7.1+ and SELinux enforcing mode by default, which is great because everyone's gonna disable it. But in theory, it's there to start off. AWS-LC integration accelerates cryptography performance while AIML workloads, 'cause there's no other workloads out there, Get direct access to accelerated drivers, including AWS Neuron support targeting customers running training and inference on AWS silicon. Previews AMI are available everywhere along with containers are in the public gallery. Feedback loops through AL27 Git repo, giving customer direct channels to influence the OS before general availability. I would expect as most people, this would remain free, as all the LLMs have been up to now— LLM-1, LLM-2, LLM-2023, and LLM-2027. [31:24] Justin: Yeah, I like that they're turning on enforcement by default, because I think that, um, more and more our interactions are via AI, and AI has a lot more patience than humans. And so, you know, if your application boots up on, you know, something that has SELinux installed and it doesn't really boot up, the, you know, the AI agent isn't going to be like, disabled SELinux! Although it might, you have It's, I have seen that actually, but not for SELinux specifically, but I've seen it circumvent sort of, I'll just turn off, I'll just turn off that feature. But I do think that it's, you know, these are good things in order to sort of guarantee your runtime and to stop, you know, some sort of executable that you're not really expecting to run in your environment. So I think, but it is sort of this challenge of like, if you're gonna launch Amazon Linux 2027, your dev environment, when you're getting started, like, it might be nice to just run a thing real quick. So I don't know, kind of go, I go back and forth on that because it's, it's fine when you know it's there and you can disable it or you can tweak it. But a lot of people don't get down to that level, especially now in the days of like containers and serverless functions. It's only grumpy sysadmins from back in the day who knew about it. [32:38] Matt Kohn: I still think it's important that it's there. It's important that, because I think it will force, if you assume most people are lazy like you and I, and we're not gonna go out of our way to turn something on, but if it's already off, I have to think about turning it off before I then do anything else. So hopefully we are less grumpy than we used to be now that AI will just do things for us and will say, Load httpd, excellent. Oh, I can't get it. Okay. There's this SELinux thing in my, in my way. Obviously I've dealt with SELinux a few hundred times in my life, but you know, at least it's there, you know? And I think it's important to kind of get those things in front of people because it's harder to disable than it is to enable things on people. So it's always easier to remove things, like to remove things than add it later on. [33:30] Justin: Amazon WorkSpaces adds support for NVIDIA Blackwell GPU instances. Crazy. So WorkSpaces applications now support graphics G7 instances with NVIDIA RTX R Pro 4500 Blackwell Server Edition GPUs, delivering up to 2.1 times better performance than the G6 instances for graphic-intensive workloads. Target use cases include CAD, CAM, 3D rendering, scientific visualization, video editing, and AI-assisted design with 32 gigabits R7 GPU memory per GPU and 2.67 times faster memory bandwidth enabling larger, more complex 3D scene streaming. 6 instance sizes are available with configurations ranging from 1 to 8 GPUs, 8 to 192 vCPUs, Hyper-V CPUs and 32 to 768 gigabytes of system memory, giving customers flexibility to match instance size to workload demands and burn money as fast as they can. Availability is currently limited to only 3 regions: US East, US East 2 in Ohio, and US West in Oregon. There are additional regions planned as capacity is required. Setup requires selecting a Graphics G7 instance when launching an image builder or creating a fleet in Workspace Applications Console. And the pricing details are available on the Amazon WorkSpaces Applications page. [35:01] Matt Kohn: You know, WorkSpaces AppStream, these like little services that I'm sure are widely used, but I feel like we don't talk about a lot, are still just some of my favorite services 'cause they solve a real business problem. And seeing them get these continuous improvements, granted it's just GPUs, but you know, people are starting to do more when they want the GPU. So I think that this is good to keep those services in the modern world and keep them up to date so that people are able to do it for maybe not just AI things, but you know, CAD and rendering and things along those lines that do require it. You know, even just general engineering, I feel like nowadays requires more and more. Compute and, and GPUs. So I think it's good to keep seeing. [35:48] Justin: I mean, it's, you know, especially with today's memory prices, right? Like, as you know, I bought a smaller Mac than I wanted to, and I play around with 3D modeling quite a bit, like designing my own parts for 3D printing and, and that. And it's, it's painful to do on this littler, uh, computer. And so like, it's, I, I kind of like the idea of maybe trying to see if I can run this in a workspace. Application and see if I can get the same, you know, 'cause it's renting a computer, but only when I need it, which would be like, you know, a couple hours and I turn it off and then I can spin it back up. Or, and so it's, it's an, I like that for that, you know, use case where I don't have to buy a giant, you know, $5,000 computer to do a thing that I might do once or twice a month. It's kind of a neat thing. [36:38] Matt Kohn: Onto our Ryan's favorite service, ECS. Amazon ECS managed, managed daemons now support non-critical daemons. ECS managed daemons now support non-critical daemons for ECS managed instances, which is the important part of this announcement, letting sidecar agents like logging metrics fail without disrupting mission-critical applications on the same instance when non-critical daemons fail, stop, or just become unhealthy. EKS keeps the container instance alive, continuing to place new applications on it and never blocks instance registration. So the app tasks launch immediately regardless of the daemon's status. Azure observability is maintained through EventBridge events on the daemon start failures and service action logs covering both critical and non-critical daemons, giving visibility into what the heck is going on without sacrificing uptime. Configuration is straightforward because everything on AWS is straightforward according to them with a simple flag in the console, CLI, SDK, CloudFormation, or I'm sure Terraform brought to you soon. This addresses a common operational trade-off where auxiliary tooling fails previously requiring production workloads to be affected and to churn new instances. But now teams can prioritize uptime because that's really what we all care about. And with our sleep, we all care about that. Over Daemon Availability where appropriate. [38:00] Justin: I think I hate this feature. So you have some sort of issue with your logging sidecar, get it, you know, and then so you, but you continue to run your container application, launching more and more instances to this instance. [38:15] Matt Kohn: Yeah. [38:16] Justin: Gaining less and less visibility into any of the application tracing or application logic there because you're not shipping logs anymore. And the only, metrics that you're getting or that you can't launch this sidecar just over and over and over. So I guess, right? Like, but if you have your app like architected well, like it, it'll, it should fail over. Like, I guess it depends on what's causing the failure of your sidecar. [38:44] Matt Kohn: I think of this like old school sysadmin. I'm gonna go back to like the days of Nagios. You got critical and you got warnings. And the way I always viewed a warning was like, I'll deal with a warning during business hours. Simplistic example, I'm running outta hard drive space. I'm at 80% of a terabyte. [39:03] Justin: I don't really care. [39:05] Matt Kohn: I got room to grow. It's, you know, I'll deal with it tomorrow. Critical, I gotta wake my ass up in the middle of the night and deal with. So to me, this is kind of where I'm thinking of like, this is sort of like a warning. Great, I don't have the logging, but do I need to fix it? And take, you know, at 2:00 AM, probably better off that we not, I'll deal with the warning during the daytime, you know? But that's kind of the way I view this where critical ones, yes, I want it to fail, we're gonna automate. So I don't know. [39:31] Justin: I guess as long as you have that operational signal so that you know there's a warning, but this looks very much like it very easily could be, you know, in the dark where you didn't know about it and then all of a sudden, you go back and you haven't had logs for months. [39:47] Matt Kohn: Well, that's where I expect to happen. Some developers can't turn it off because logging keeps crashing due to some random reason. They deploy to production and they forgot that they turned on this flag. That's what I kind of expect to happen. But in theory, you set up the EventBridge to send notifications, you know, and you monitor it through your warning channels or your less criticals. So again, could be used for good. Ryan and I will use it for evil. [40:13] Justin: Got it. Yes. All right. Amazon GuardDuty adds optional threat detection rules. GuardDuty has added 35 opt-in custom detection rules covering CloudTrail management events, producing 26 finding types mapped to the 10 MITRE ATT&CK tactics without requiring customers to manage log ingestion, normalization, or storage. The rules address context-dependent threat indicators such as external IAM sharing, disabling flow logs, or MFA-less sign-ins, letting customers enable detections only for activity that is genuinely anomalous in their environment rather than routine. There is a dry run mode and allows teams to test detection efficacy before enforcing rules live, reducing the risk of alert fatigue and false positives during rollout. Available now in all AWS commercial regions and GovCloud with access via the GuardDuty console, API, and pricing, of course, follows the existing GuardDuty model. [41:12] Matt Kohn: I like new rules. I like that they're optional. I still question whether MFA sign-ins are ever anomalous, because I feel like they should just be disabled. I feel like that's something I've been yelling at people for 10+ years. Set up MFA everywhere. [41:29] Justin: Yeah, I mean, it's the trust but verify. I would disable MFA. Or disable non-MFA logins, but then I also want this check to be like, oh, somehow someone got in without it and now it's a security finding in my console, which is what I like, you know, like, and it should be an anomaly. I shouldn't see this ever if I've set it up right. And so it, you know, it'd be something where it's like, 'cause I've had this happen where, you know, you set up your rules and your device trust and all these things and then there's some edge case somehow that allows a workload in. Where, you know, or if you've, especially if you got a specific type of MFA device and you want to enforce that specific one, it's, you know, I like the fact that you have this sort of, you know, security notification that you can go investigate based on that. [42:16] Matt Kohn: Speaking of new security features, AWS API Gateway now supports mutual TLS for backend integrations. API Gateway REST APIs can now present real ACM-issued certificates during TLS handshakes with backend integrations, replacing the previous self-signed certificate approach. This closes the mutual— this closes the loop on mutual TLS since inbound client-to-API mTLS was already supported. Certificates can be imported from existing PKI or issued via the AWS Private Certificate Authority. Is it still $400 just to turn it on? [42:56] Justin: I have no idea. [42:57] Matt Kohn: I think when it launched it was. I don't think I've ever gotten that out of my head. It was, yeah. Giving customers the flexibility depending on their existing certificate authority relationships. AWS ACM handles renewals and re-imports automatically with API integration propagating updates without redeployment or downtime. Targets This is targeted to regulated industries as security is more important and mTLS is becoming a bigger and bigger thing. It's available in all regions now for whatever RESTful APIs that you have supported. It is a free feature, though there are costs for ACM and obviously the API Gateway that you are using. [43:41] Justin: Yeah, I mean, if I am going to run an application that does mutual TLS, I'm I'm only willing to do it if I'm using something like, you know, Certificate Manager or managed service that's handling the certificates on my behalf, just because it's so painful to coordinate. [43:56] Matt Kohn: I was gonna say, you've never run your own PKI infrastructure with a primary and an offline one and where you put that secondary key and you hate your life a little bit. Yeah. Just remembering the start. [44:07] Justin: Never again. Never again. Like you were talking about disk space error and I'm like, oh, I'm never dealing with that again. Like if I have a disk space error, that thing better die and stand back up. Like, uh, outages during that kind of stuff. Like I'm just, you know, like it's so difficult to keep track and these things expire and especially as, you know, certificates and rotations becoming much more short-lived. Um, I do like this, right? Uh, as long as it has that full automation end-to-end. Point where you can sort of just turn it on and forget it. And then, but your, your connections are being validated both ways. And so your client knows that they're talking to the right, the right place and can validate that this is the right client as well. So it's good. All right. Moving on to GCP. Getting started with the Mantis harness to find and fix bugs, which is a Google Cloud Google has open-sourced Mantis, an agentic bug-finding and patching harness used internally to automate vulnerability discovery, triage, and reproduction and fixing at scale. The tool addresses a known weakness in AI code scanning. Standard approaches also often produce hallucinated bugs with true positive rates under 7%. Mantis improves accuracy by combining critic and review agents with sandbox reproduction to validate findings before flagging them. A hierarchical security summary tree condenses file-level detail into directory and root-level summaries, cutting token overhead by over 85% while retaining the architectural context. Mantis learns from a repository's own commit history and builds to build architectural and threat model documentation automatically, even with none, even when there was none that previously existed. And the setup is as simple as cloning the repo and prompting an agent to use the framework against its target codebase. Google recommends pairing Mantis with human-curated context, for example, defining which bug classes are out of scope, and a dedicated sandbox with clear vulnerability acceptance criteria, and a companion Mantis advice skill helps coding agents write more secure code going forward. [46:15] Matt Kohn: I really like the dedicated sandbox. [46:19] Justin: I was really hoping that, because like the fact that this is getting started with the Mantis harness, like Like, you know, in AI everything's a vague term. And so I still define harness as like the agent runtime as well as all the instructions, but not everyone does. They just say the harness is the agent instructions plus the model. You know, and it's like, I like to see more of these harnesses with a more sort of prescriptive sort of runtime approach, including that sandbox, like that isolation. I want something that's a little bit more turnkey. I think this is neat for sure, and I, I definitely will take anything I can to, to add to my agent instruction already so that I don't have to write and contribute more to my heap of Markdown. But, you know, like it is tricky to run these things in runtime when giving it the access that it needs and giving it the execution environment that's safe and secure and isolated. 'Cause we've learned that's a major issue now. And I was sort of hoping this solved it, but I don't believe it does. [47:24] Matt Kohn: Yeah, I mean, I kind of feel like we are at the point still of like, here's the software, you figure out how to run it, versus here's the Docker container. Like, we're getting there to that end state of don't shit my— don't tell me it runs on my computer. You know, okay, great. I'll ship your computer to production, you know, but like give us the entire thing. But I don't think AI is even there yet. Like it's too, when you talk about harnesses and agents and all these things, like there's still the sandbox that's missing. There's all these other pieces that are so missing that you have to still build. I think eventually we'll get there. I just think, don't think you're quite there yet. Yeah. [48:10] Justin: I mean, it's, it's just, it's not prescriptive enough. And so it's like, I think that it's. We're getting there, but I, I, you know, like my problem is really like we're calling, you know, a set of Markdown tools and skills and stuff, the harness, and not everyone really sort of agrees with the definitions. And that's sort of my problem is that, you know, like a lot of times talking harness, it was more about the execution of it, like talking about using Claude code and that and VS Code and, and those things that are executing those agent loops. Like for me, that is the hardest. It includes those, you know, the things that are loaded up at agent runtime, but I, you know, we don't really have the greatest sort of software solutions, which is what, you know, like I think Open Claw tried to solve and these things which ran into problems. So I'd like to see it sort of more of like, this is the Docker container and this is how you use it. Which it's not as prescriptive as I'd like, but you can run it in a container and it does use a container to set up sort of sandbox execution when you've sort of configured it for it, but you have, it's not. So it, it, it executes its own sort of sandboxing within the instructions, but then it itself is still running on whatever bare, you know, system you have it on, just sort of silly. [49:39] Matt Kohn: Yeah, I'm gonna point it at some random AI slop I've generated and see what it does. Maybe we'll point it at Bull Bot and just see how much it rips apart Bull Bot. We'll have some fun with it. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber. Google shipped its third Flash release in 6 weeks. So I really hope that your CI/CD pipelines have all, and all your testing is set up for this. With 3.8 Flash and a specialized 3.8 Flash Cyber, both priced at $0.75 million per input tokens and $3.75 per million output tokens, matching 3.7 pricing. 3.8 Flash targets long horizon, long horizon code and agentic tasks. Outperforming longer frontier models such as DeepSWE 1.1 and scoring 54.9 on the HLE verified. It achieves this by using more reasoning steps and tool calls, so token usage can increase at higher effort settings. 3.8 FlashCyber is restricted to trusted defenders through the new Fairwinds program focused on vulnerability discovery and automation patching rather than offensive capability. Can you guys see the trend of today? Security tools and AI. Google cites internal validation. Chrome security saw 2.6x more correct patches versus larger commercial models. Wiz, which clearly is now part of Google, reports 7.5 to 9.7% higher recalls at a 2.3 to 5.2x lower cost. And Google Cloud Vulnerability Research team found a critical vulnerability in under 2 hours using the model, which makes me question Google Cloud, but we'll bypass that. AccessPoints span the developers and enterprise stack. It's available in AI Studio, Android Studio, Google Antigravity Enterprise Gemini for 3.8 Flash, et cetera, et cetera. Cyber requires applications through the Fairwind program. Look, another Flash model. Flash is here. That should have been our show notes. Show title. [52:06] Justin: I've been under the impression that the Flash models were smaller set of Gemini models for loading onto like, you know, smaller hardware, but I don't think that's correct. I think Flash is a name, cuz it, I was just looking at the, there's Flashlight. And so I've been under it, like, I guess it's like more of like the Sol Terra model where there's Flash and Flashlight and it's kind of interesting there. Uh, so sorry, I'm learning on the fly. You guys are having a— [52:34] Matt Kohn: I thought the Light was like, you know, the lower, lower end model, you know, kind of like that. [52:39] Justin: But maybe I thought Flash was, um, but I guess there's not, there's There's flashlight. Oh, I get it. Flashlight. [52:45] Matt Kohn: Uh-huh. [52:48] Justin: Because I was sort of surprised to see that, you know, there's a cyber model being, you know, but with a smaller token context window, right? Like size. But I guess that's not really what it is. It still looks like a fairly, you know, big multimodal model with, you know, a million token context. So. Yeah, I guess everyone's getting on the cybersecurity bandwagon. Um, I guess that's sort of, I wonder if that's just, you know, the market reaction to Anthropic's Mythos sort of thing. And then they're, they're also sort of following, they were also waiting maybe after OpenAI paused their stuff, 'cause it all sort of seems to be coming, you know, back. I don't know. We'll see. I do think it's neat. Like, I do think it's gonna be a great use of AI is to, to sort of protect this, but it's also, I'm hoping that the safeguards and the protections are good enough, and I'm hoping that people can't use this for evil or it doesn't get out, that the controls for who gets access to these models through the different programs, whether it be Fairwind or Glassway, hopefully those are secure enough. We'll see. [54:01] Matt Kohn: We shall see. [54:02] Justin: All right, moving on to a story that we're only talking about just to appease our, our children, or at least, you know, Justin and I. Matt's children are a little small, little too young yet, but you know, give them time, they'll be part of the YouTube generation soon enough. And, and, uh, but Google is expanding its multi-year relationship with MrBeast and Beast Industries beyond its YouTube sponsorship, and it is expanding that relationship into Gemini and Google Health integrations. Marketing Gemini through a high-profile creator with 500 million subscribers. MrBeast's September 5th video will feature Gemini being used to help his team navigate survival challenges in the jungle, deserts, and Arctic environments, positioning Gemini as a tool for real-time hazard identification and decision-making. The partnership includes a Gemini ad campaign spot showing MrBeast using the app to coordinate logistics in his large-scale video productions, an example use case for AI-assisted project planning. Google Health and Fitbit are also a big part of the deal, with Fitbit Ace integrated into an upcoming MrBeast challenge, tying consumer wellness hardware into the broader promotional push. For our listeners, this is just a consumer marketing story that we are shamelessly talking about just because we thought it was funny, and so our kids like us. [55:19] Matt Kohn: I mean, we talk about it all the time, how we use AI in software development and different things in our life. But it's interesting to me to see how other people are using AI and, you know, what people are doing, you know, to leverage it in different ways. 'Cause I feel like that's really the way we're learning. So like, I know I personally used it to read terms and conditions of credit cards to figure out some nuanced details because I wasn't reading the 57 pages, but AI could read that in 15 minutes. And 15 seconds to tell me all the details and questions I needed to know. [55:53] Justin: I'll read the 10 bullet points if you summarize in hopefully 10. [55:57] Matt Kohn: 10 might be too many for you, Ryan. Yeah. You know, but like, I was just signing up. I was like, well, if I sign up now and I closed it because I— there was reasons I closed it, like, will I get the signup bonus? And like, I was like, here's my situation, go figure it out. So I think it's interesting to see. Yes, I understand it's ads and marketing and everything, but I think it's still interesting to see how other people leverage AI in day-to-day versus you do, because there's things that, you know, we talk about before the show, after the show, in the host channel, in the general channel. It's like people say things like, I just never thought to do it in such that way. And like, it just like, you know, blows a little portion of my brain out. Like, oh, I can now do all these other things that I've never thought of doing. [56:41] Justin: It is, uh, fascinating. And I, and I talk to a lot of, you know, like my non-technical friends and they're either not using AI or they're sort of dabbling and don't understand it. And so I, I agree, this is sort of a neat little example for me just in terms of, uh, like, oh, you know, planning and identification, like especially the planning. Like I know some people use it a little bit for like itinerary and planning trips, but I think it's, I think it's really surface level in a lot of areas, um, for a lot of people outside of the technical technical fields. And so I do, I agree that this is sort of a neat thing and I'm always looking for more examples of how to use it like outside of like sort of my day job and technical work. [57:22] Matt Kohn: Home repair, in case you haven't used it there, it's fantastic. [57:27] Justin: What, and it loaded up into a RoMod? Like I still think I have to replace the sheetrock, don't I? [57:31] Matt Kohn: Yeah, okay, let me correct that. How to fix things at home. [57:38] Justin: Okay. [57:39] Matt Kohn: Or like finding stupid little parts. Like we have this like inlay on one of our windows in our house that's like this old inlay and all the little pins fell out. I was like, here's a picture of the door. Here's the part. I gave it like 4 pictures. I gave it, I know it's an Andersen door and it told me the part model in 3 seconds. [57:57] Justin: Nice. [57:57] Matt Kohn: And I was able to buy it. Like little things like that are like, how do I like do this thing? I don't remember what it was. I was doing something the other yesterday. and I was like, here, it's like, here's the 6 instructions how to do it. And I followed the instructions. So like, to me, sometimes it's just a faster Google search that like finds exactly what you're looking for with your model and context and everything. [58:17] Justin: Yeah. I guess I did use it for, for planning out like sort of a solar, solar build out on my RV and it's, it made it a lot easier to sort of understand the context of like solar capacity and recharge rates. 'cause it can be a little complicated when you start talking about things that don't have 100% efficiency and, and things with variable sort of electric load. [58:38] Matt Kohn: Yeah. [58:39] Justin: So yeah, you're right. That is sort of a, you know, you can get more complex conversational sort of context or just bouncing ideas so that you're not like, I wanna buy this, these things specifically. And it's like, oh no, don't do that. [58:51] Matt Kohn: Yeah. I was talking with my father-in-law this weekend and they're having a problem with one of the car, with one of their cars., and it just like randomly like goes into park when you're at like approaching, like, you know, stopping at a light. It just turn, like, we'll flip to park. So you get this like switch, but it's completely inconsistent. Yeah. [59:09] Justin: No. [59:09] Matt Kohn: So like we've had like, like I was like, why don't we ask? Yeah. Like we had that conversation with Claude and it's like, hey, check this thing. And we're like, well, the first thing it's like, it's a $20 part and a 4-second fix. [59:19] Justin: So like, nice. [59:21] Matt Kohn: It might be wrong, but it's still such a cheap part. [59:24] Justin: Yeah. [59:25] Matt Kohn: And a 15-minute fix of just crawling on the dashboard, changing one thing, like, that's worth it to try if it happens to be correct. Oh yeah. He went to 3 mechanics, none of which who told him that this could be the problem. And it was like, this is the potentially the highest thing problems. Like, it's interesting, like, given our conversation now, like, I never would've thought to use it probably for solar relay cuz we're talking about solar on our house. [59:48] Justin: Mm-hmm. Yeah. And you can feed it all kinds of information, right? Like your electric bills and how much data, as you're saying. And yeah, and it can probably search tons of public information about sunlight and, and, you know, information about different solar panels and stuff. So it's, yeah, I mean, that's kind of neat. [60:09] Matt Kohn: In other good ways that we use AI, whether Next 3, our most advanced global AI model for weather has been released. WeatherNext 3 shifts training data from lagged NWP simulations to live geostationary satellite mosaics and sparse weather station observations, enabling hourly forecasts up to 5-kilometer resolution, roughly 5 times sharper than the WeatherNext 2, which had a 25-kilometer and a 3-hour cycle. Well, that sounds very impressive. I only understand a very high level of this. Precipitation forecasting shows measurable accurate gains with CRPS improvements. I don't know what CRPS is, so I hope you do. Up to 60% against NASA IMERG data with a 30% MRMS and 10% against rain gauge measurements. I know that for early, for early lead times addressing longstanding weak points in AI weather models. [61:10] Justin: Prometheus. So I don't know the specifics, 'cause I'm not a meteorologist, but I did watch sort of a, it was kind of, I think it was, I don't know if it was related to this or not, 'cause this is sort of a new model from Google DeepMind and Google AI research that they're releasing out there. It's not very clear from the press release that this is a new, it's a model specific for that. The thing I was watching was sort of a, you know, we've had forecasts, weather for a while now, but it is very different to do sort of a physics-based machine learning data crunching sort of forecast versus what AI does. It doesn't replace the old way of doing forecasts, but it's a new way to do it on top of that. Those forecasts where AI is going to take not necessarily just all the raw data, but it's trained on those datas and then it can make predictions based off of the incoming inputs. And it can do that at a much faster rate than you can do like a full run of a forecast where it's gotta go to like some supercomputer and it crunches out all the physics numbers from the readings and stuff to determine like the pressure differences. And I'm doing a sort of bad job rehashing what I saw, but I thought it was very interesting, the two different ways of getting that data. And then, you know, what, you know, AI allows you to do, it, it'll, it might be a little less specific, but you can run it 1,000 times, you know, uh, very quickly comparatively to like the one, one you can do it for the physics model. So it's kind of neat. And I do think that it's, you know, I think AI and the model training that you can do with these types of things, is you can refresh it a lot faster and retrain the model, which I do think is, you know, a heavy lift. But, you know, it'd be kind of neat if we can get to the point where we really can predict weather more specifically. You know, right now we know, like, you know, there's a lot of hurricanes right now, um, circling the, you know, Hawaiian Islands. And so there's— we think it's going to go over here and it's gonna, it's gonna graze the island, but if it, if having a more precise prediction would be great for those people who are just probably living on edge right now, hoping that they don't get crushed by the storm. [63:32] Matt Kohn: I mean, that's kind of where my brain went with this was hurricane and spaghetti models. And if you can tell from Florida, from using the term spaghetti models, I don't think people outside of that, it's that model that shows all the lines of where it goes and, you know, It also bothers me that in the article it calls it cyclone heading, but it's a hurricane in the Atlantic and it's cyclone elsewhere. But we'll bypass that nuance because it shows the state of Florida in there. But like, if you can get that down, that would help a lot, you know, let alone time, because time is the biggest killer of hurricanes. If a hurricane hits a high tide, it does a lot more damage than hits at low tide. And this is speaking from someone that's lived through many hurricanes. Being from South Florida. And, you know, two hit 7 miles apart from each other, 5 miles north of my house growing up, you know, but they all hit at low tide, so they weren't as big. And I was not in the northeast quadrant, which, if you don't know anything about hurricanes, is the moat— is the strongest portion. So I luckily was spared because of being low tide and not being in that quadrant. But a 15-mile shift and my house would have been destroyed when I was a kid. So If they can get that down, that also would help not just people prepare, but also people leaving areas because people leave and then get hit because they moved. And also, you know, help FEMA and other companies, you know, other, you know, gas companies, electric companies, everyone really help them prep for the hurricane and know where to go right when it's there. So people are standing by outside storms. I have a little bit of hurricane scar tissue in my life, in case you can't tell. You take the boy out of Florida, but apparently can't take hurricanes out of me. Yeah, I still track hurricanes even though I don't live there. I have family that live there, my parents, my sister, you know, I was the third generation. My grandmother was 100 and moved there when she was, when she was 21. I was born in the same hospital as my dad. Hurricanes are ingrained in me at this point in my life. Yeah. So yeah, on to Azure. [65:34] Justin: Oh, if we have to. [65:35] Matt Kohn: There's only one. [65:36] Justin: There's only one. Yeah. So Windows Server 2025 on AKS is finally generally available. In the year 2026, we get Windows 2025, giving a path forward as older Windows Server versions approach the end of life support. Key improvements include stable ABI. What's that? ABI? Is that real or is that a typo? Generation 2 VM is the default containerd 2.0 runtime and FIPS compliance enabled by default for regulator workloads. This targets enterprise running Windows-based containerized workloads, you poor suckers, who need to modernize their AKS clusters ahead of the Windows Server 2025 lifecycle deadlines. The FIPS by default setting is notable for organizations in government, finance, or healthcare that require compliance with federal cryptographic standards without additional configuration. No specific pricing details are provided in the announcement, but standard compute and Windows Server license costs likely apply based off of your node pool configuration. [66:44] Matt Kohn: Hooray! I tried to do the reaction, but I couldn't find which one does the audio versus just the picture. [66:49] Justin: I was trying to do the blow horn. [66:52] Matt Kohn: Better. Yay. [66:55] Justin: I mean, yeah, I'm still stuck from Windows containers. Yeah. [66:58] Matt Kohn: Oh my God. I hear they've gotten better, but they're just hard. We'll go with that. [67:08] Justin: Yeah. It's, it's definitely not built for it. Right. Like, I feel like they've had to come, come at cramming Windows into this, you know, square hole, trying to get it to work. Where it was, you know, containerized and, and, and that sort of isolation was, has been in like Linux for many decades at this point. So, yeah. [67:30] Matt Kohn: Well, then they also like shimmed Hyper-V into be the container, you know, the way you ran a container. Like it's all sorts of shim here, shim there, shim there. I think it's gotten better since we've done it, what, 2016, 2019, but I still would not wish this upon anyone. So we have no Oracle stories, but we do have one emerging cloud story. Introducing context-aware vulnerability discovery and remediation— feels like a theme of the today— with Cloudflare Managed Defense and OpenAI Daybreak models. So this is kind of piecing together a couple stories from earlier. Cloudflare is launching the Vulnerability Discovery and Remediation by Invite Only that pairs the OpenAI Daybreak models, including GPT-5.6 Cyber, with Cloudflare network data to prioritize which vulnerabilities matter most than just your compliance team listing all your vulnerabilities and telling you to fix them all tomorrow. The key differentiator is production context. The system context references code vulnerabilities with actual traffic, active routes, and existing WAF rules to determine real-world exposure and addressing common problems. Security flags thousands of things with no way to actually write them. The architecture keeps humans in the loop. Models can propose code patches and WAF rules but cannot directly implement them, which is good so it doesn't just shut down everything. All proposed pass through validation checks and customers review before being deployed. Technical workflows use multi-agent pipelines reconnaissance, hunting, validation agents that map production routes to source code with models inference happening on the OpenAI servers via CloudFront AGI gateway, or sorry, AI gateway rather than at the edge. This builds on CloudFront internal vulnerability hardness previously discussed with your build your own vulnerability post that they've talked about. And I believe we talked about and extend extends the fleet scanning capabilities to customer codebase, signaling a broader trend of combining LLM code analysis with infrastructure-level telemetry for security prioritization. [69:41] Justin: Yeah, I mean, this was the first thing I started thinking about the minute we started getting reports of like AI vulnerability, you know, AI is stitching together many different vulnerabilities to, for exploits. I started thinking about doing it, doing this for, for workloads, just because for vulnerability management has always been a problem tracking sort of meantime to resolution. You know, when you've got a page of like 40,000 software detections of vulnerabilities, but you know, because of the contextual situation of your environment, that there might be one or two that's actually sort of a problem, or at least a critical problem. And when you consider sort of like just updating a vulnerability has a cost, right? The updating it alongside your code feature base and running it through the SLDC to make sure there's no impacts from changing libraries, changing dependency models is, is, is a thing and it has a real cost. And so it's working with dev teams from a security perspective and understanding that cost and then also trying to get these updates. In their general flow has been key. And so this type of prioritization based off of the actual context is awesome. Um, I, I'm really curious to see if this is running off of like their larger base of network data or if the, if they're tailoring this to your specific sort of CloudFront configurations. I don't know. [71:06] Matt Kohn: I assume they're tailoring it because they talk about your paths and what's happening with your stuff. I assume there's a base of everything, but this all feels fairly tailored. I mean, that to me is really going to be the value is, you know, when I'm sure you were doing it and I've done some of it where I take some network traffic and some code and try to piece it together, but this does it in real time consistently, which is really nice. And I'm doing it more for bugs rather than security, but it's definitely a good spin on it, which is WAF rule automation. And things along those lines that can start to put in the WAF rules in place until you have time to fix your code base with something like that. So I'm imagining something like Log4j again out there where Log4j went out, it would tell you what's being attacked and you would press a button that does it. Now all the cloud vendor, all the vendors when Log4j came out, all just had WAF rules available. Yeah. [72:02] Justin: Added the WAF rule classification. [72:03] Matt Kohn: Yeah. [72:03] Justin: Where you just, Now it's checkbox. [72:05] Matt Kohn: Now it's like, do you, do you need it or not? And where do you need it? You can kind of pinpoint a little bit more with something like this, which I think is pretty cool. Yeah. [72:14] Justin: I mean, thinking about stuff you just can protect with the WAF is, you know, it's a definitely, you don't want to totally lean on it, but it is a nice perimeter thing. You know, Log4j is a really good example of something where you can, you could put in a single rule and get a very wide protection. SQL injection is another one where it's like, if you can detect these things into the payloads, it's, uh, it's pretty good to detect it at that point and then get a very broad protection. Yeah. But also you should sanitize your inputs. [72:45] Matt Kohn: Details. [72:46] Justin: I know. [72:47] Matt Kohn: And sign your codebase. Also, AWS has a service, AWS Signer. [72:53] Justin: So anyway. [72:56] Matt Kohn: Ryan, we survived. [72:58] Justin: We made it. [72:59] Matt Kohn: We weren't too cynical, I think. [73:01] Justin: Oh, I doubt that. We'll have to listen to playback and maybe we can edit this into some positive and coherent conversation. We'll see. [73:10] Matt Kohn: We gave Elliot some work to do this week. [73:12] Justin: With some challenge. Yeah. [73:14] Matt Kohn: Yeah. [73:14] Justin: All right. Bye everybody. [73:16] Matt Kohn: Bye everyone. Another week of cloud news wrapped up. Bolt will collect the news. Justin will get the Jonathan will write some code, Ryan will watch the perimeter, and Matt will reluctantly watch Azure. Till next week for AI, Amazon, Google Cloud, and Azure, and hey, maybe even Oracle, who knows? Check out thecloudpod.net for our newsletter, join our Slack, message us on socials, or leave a review.