# 363: SQS: 20 Years of waiting in line Duration: 79 minutes Speakers: Matt, Justin, Matt Kohn Date: 2026-07-22 ## Transcript [00:07] Matt: Welcome to The Cloud Pod, where the forecast is always cloudy. We talk weekly about all things AWS, GCP, and Azure. [00:14] Justin: We are your hosts, Justin, Jonathan, Ryan, and Matthew. [00:18] Matt: Episode 363, recorded for July 14th, 2026. 20 years of waiting in line. Good evening, Ryan and Matt. How you guys doing? [00:28] Justin: Doing good. Nice. [00:30] Matt Kohn: Yeah, I'm awake and I'm here, so I call that a win. [00:33] Matt: Yeah, I, uh, I got home from Tennessee very late last night to bring the humidity apparently from Tennessee with me to my house where it was 100 degrees and humid. So it's been kind of miserable today here at the, uh, the homestead. But, uh, hopefully it cools off pretty quickly because I don't know if I can do that much humidity that long. [00:49] Matt Kohn: Yeah, and I walked out of the house right before we record to take the dog for a walk, and it's currently 92 degrees at 9 PM. I was like, that's great. It's gonna be a good night here. [00:59] Justin: Humidity really holds it in. Like I'm, you know, just fresh back from kind of the Midwest. They don't like being called Midwest, but it's in Ohio and the, the humidity just really seems to hold the heat. So like there's no cooling off like you get sort of in California at least. It's gross. I don't know how people say— [01:16] Matt Kohn: Born and raised in Florida. It's 100% humidity or nothing. [01:20] Justin: Yeah. [01:21] Matt: That's why I don't like spending a lot of time in Florida. Yeah. It's miserable. [01:23] Matt Kohn: Yep. If you're not walking outside and sweating, you're not doing it right. [01:28] Justin: Yes. [01:28] Matt: Uh, all right, well, we've got a bunch of cloud news as always here. So first up, uh, former GitHub CEO Thomas Dumke has launched Entyre, a startup building a distributed Git network aimed at reducing load on centralized hosting caused by AI coding agents. Company's apparently raised $60 million seed round at a $300 million valuation. I mean, for not having a product, that's pretty darn impressive. Uh, I mean, basically, if you've heard about GitHub's woes, and we've talked about here on the show, they're saying that basically the exponential growth of code is part of the reason why GitHub has not been stable. I think it's a convenient excuse, just like it's a convenient excuse for layoffs, but, uh, you know, that's where we're at. Apparently the core idea of this though is mirroring GitHub repos across regions, US, Europe, Australia, so agents can pull from nearby mirrors instead of hitting a single central repo, which addresses a rate limiting and latency issue that comes with high volume automated cloning and pushing. Reported internal benchmarks include about 570,000 clones per hour on a single repo and 586 pushes per second. Though these are self-reported and not yet independently verified, Entire says it plans to open source the Git backend and benchmarking tools for third-party validations. Beyond distribution, Entire is building a semantic layer on top of Git history, capturing agent prompts, reasoning, and tool calls, with features like Entire blame tracing, AI-generated code back to originating prompts, and Entire review support, multi-agent code reviews. This treats AI agent traffic as a distinct infrastructure problem separate from human developer workflows. So, I mean, it's Git for AI. That's basically how you could read it. [02:50] Justin: Kinda. Or it's a, it's memcached for Git, you know, like it's, it's, it's kind of just that, right? [02:56] Matt: Like geo-replicated Redis for, yeah. [02:59] Justin: Mm-hmm. Like the numbers sort of bear out some of the complaints that GitHub had in terms of like the load. Like, 'cause you think AI writes some bloated code. Yes. And then all the agents are working much at a much faster pace than, than humans. But I do think it's also just kind of shortcoming in, in capacity forecasting for Git, right? Sounds like, you know, the former CEO has a technology solve, which it's pretty smart. Get fired for having all these issues and then solve the problem anyway. [03:26] Matt: Like, okay. And then get bought by them probably is your plan, right? [03:29] Justin: Most likely. [03:30] Matt: Come back in as the shining hero. [03:33] Matt Kohn: I do like the idea of splitting out the Git, the AI workflow and It's really interesting to kind of see how, or it will be interesting to see how they actually like track it all the way back. So essentially like to get blamed for your prompt, to educate people because, you know, I'll have a conversation, build something, kind of go from there, but then I never know it went south and I'm like, okay, what prompt did I make? Or what did I do? Or what, you know, plan, thing in the plan did I miss that made it go south when I wanted it to go north. So if you could track that back, it would be an interesting, you know, way to self-improve upon your prompting. [04:15] Justin: So I don't think this will separate out the traffic in any way to give that visibility. It really is just, because it's, you define the mirrors as targets, it sounds like. [04:25] Matt: I mean, how does it, I mean, the idea that it could, you know, do a blame back to the prompt, the originating prompt is interesting. Like, how exactly is not in a GitHub repo today? That I'm aware of. [04:36] Matt Kohn: Not that I'm aware of either, unless if they're gonna require you to use like, you know, originally GitHub Actions or CLI, some proprietary, right? [04:43] Matt: Yeah. [04:43] Matt Kohn: Something like that in order to leverage it. [04:46] Justin: Hmm. [04:47] Matt: Interesting. I don't know. We'll just keep an eye on it. We'll see if, uh, you know, $300 million valuation already, so they gotta do something. [04:53] Justin: Yeah. It's, I mean, it's amazing considering how nothing this is, like it's just an announcement of an intent to do a thing. [05:00] Matt: Like, mm-hmm. I mean, I, I would definitely invest in the former founder of GitHub. He's, you know, he has probably good ideas, and I imagine that you're funding the team more than you're funding the idea at this point. [05:12] Matt Kohn: You know, every time we talk about something like this with these, with these massive evaluations, I always think of ai.com, and every time I go there, still in beta. [05:21] Matt: Yeah, I don't know if anything's ever going to happen to ai.com. [05:24] Justin: That is really— [05:24] Matt: I don't think it's either. [05:25] Justin: They just spent a bunch of money on a football ad and it's just nothing. Wow. Yeah. [05:30] Matt: I'm surprised like Wired Magazine or, or The Verge or somebody hasn't tracked down what the hell happened with that whole thing. 'Cause like someone out there spent a lot of money on that ad and yeah, it's just like, where's my Fyre Festival documentary on that? [05:42] Justin: Yeah. [05:42] Matt: Right. That's what I wanna know. Well, Satya Nadella, uh, wrote a Twitter post, uh, over the week. I don't know why he decided to write these all on Twitter now. Like, uh, it's the least popular place I want to go read your blog thoughts. It's on Twitter, but okay, whatever. Uh, basically Microsoft CEO Satya Nadella coined the reverse information paradox, which he says is AI flipping Kenneth Arrow's classic info paradox. Instead of sellers giving away value before being paid, buyers now have to feed proprietary knowledge into a model just to make it useful, paying twice, once in dollars and once in your intellectual property. Nadella says every prompt correction and eval is exhaust that trains the model's prior futures intelligence, a one-way flow where the vendor learns about your business and you learn nothing about their model. Without a fix, economic value drifts to whoever owns the learning infrastructure, the model provider, not the companies actually generating the knowledge. He quotes Palantir's Alex Karp on customers wanting to own their compute, models, data, and alpha, not quietly hand it over. Vidalia says enterprises need a hard trust boundary, a line across which nothing, not even usage exhaust, crosses without consent. So data, traces, evals, and tuned weights stay owned by the firm. His 5C framework for this is control your own evals, traces, feedbacks, Built-in tenant compatibility to models. Keep choice by decoupling orchestration from any single model. Use decoupling to manage costs and compound it into a continuous learning loop. Using a model shouldn't require giving up the knowledge that makes your company unique. Which I think is fine. I get what he's saying, but I also, you are selling one of the largest LLM models that's literally sucking up all this knowledge and doing exactly what you just complained about. So what's your solution, Satya? That's what I want to know. [07:21] Justin: And if he hates this, I really hate to see his reaction when he finds out about advertising tracking cookies. [07:30] Matt: Touché, touché. [07:32] Matt Kohn: I mean, isn't the premise of, you know, running it all inside of Azure or AWS that they're not training on your data? [07:40] Matt: I mean, I think that's the assumption, but like, that's the data you're actually feeding into the prompts, right? Like, my data is not going to be used to train your models is the claim, right? And that's the contract that we have in place if you're an enterprise buying Claude or you're buying, you know, Azure, OpenAI, whatever. But that being said, even without having the prompt data, there's a ton of data that's there considered theirs intellectually around how it thought about your data. It doesn't have the data, but it knows how it thought about your data. So I assume that you could probably infer and, you know, infer some of the data model without even having the data to train your model because you have all this output and thinking logic around it. And I think that's what he's trying to point out is that that's the risk. And no one really has a solution for them unless you're using your own GPUs and you're running your own models on those GPUs, which is what maybe he's trying to say Azure can do for you. But again, it wasn't clear because he didn't call out how Azure solves this for you, which is what I would hope someone would say. [08:34] Justin: But, uh, yeah, and I, I don't quite know if, like, if what he's saying is all— is more valuable, like, than any, like, any kind of tracking of user behavior and accessing an application or a site. Like, I don't quite— like, it is an interesting look. There's, there's data that's not your data, and that's probably not very clearly defined, and it's very easy for a company like Anthropic or OpenAI to be like, well, that's not your data, that's, that's this data of usage, and so therefore I can use that. And you can probably get a lot of data out of those prompts and and the responses, probably just the prompts, um, and then use that, which I think is where he's getting it. And then, but I'm like, I'm not sure that the value is there for me to be concerned other than the normal privacy. Like, I don't, I don't, I don't use AI as a therapist. I don't use it, you know, to, to crunch my numbers for finance, like, kind of thing. So like, because of privacy concerns, I don't know. [09:39] Matt: Yeah, there's definitely risks. I mean, I also, you know, in many lawsuits apparently the AI will tell you to go kill yourself. So be careful using it for therapy. [09:48] Justin: Yeah, do not. Yeah. [09:50] Matt: Don't pass go on that. Well, yeah. So I'm curious, you know, if this is him starting a series of blogs then leads to some announcement or something that solves this problem or if he's just, you know, putting it out, speaking into the world. [10:04] Justin: It's weird. Yeah. Cause I'm like, does he see something I don't see? Like, I, like, it kind of feels like that. Cause you know, like he's been. Smart. He's definitely done wonders to, to Windows and Microsoft in general. [10:14] Matt: So yeah, I mean, he's, he's exited us from the Ballmer era quite successfully. Like, people respect Microsoft somewhat now, maybe not as much as they'd like to be respected, but still better than they were during the Ballmer era. And so, yeah, they're not trying to make a phone anymore, which is great. Yeah, apparently they're not trying to make a new version of Linux. I don't know about that part, but, you know, here we are. So Uh, all right, let's move on to AI is how machine learning makes money. And for this week, you know, if you're curious about some of these models, there's a great blog post that we're talking about here from Try AI. They did a build-off test of Groq 4.5, GPT-5.5, and Claude Opus 4.8 and Claude Table 5 on identical one-shot prompts to build a 3D Rubik's Cube, a particle gravity sandbox, a breakout game, a self-contained HTML files, and then measure the latency and cost through a unified testing harness. Uh, and some of their findings were quite interesting. So Groq 4.5 differentiated on performance metrics. Sub-half-second time to first token, roughly 110 tokens a second throughput, about double that of the other models, and the lowest cost per reply, though it required one retry to render the Rubik's Cube correctly on the hardest task. Claude Opus 4.8 and PaLM-5 were the only models to successfully render the 3D Rubik's Cube with the correct colors and animated solving on the first attempt, positioning them as a more reliable choice for complex, stateful coding tasks at the cost of higher latency and price. And GPT-5.5 produced the most visually appealing particle gravity sandbox and had the fastest response times on short answers, but failed to render a complete Rubik's Cube. Yeah, it's pretty sad, the Rubik's Cube, if you look at it. All of our models successfully build a playable breakout game on the first try with score tracking and lives, indicating that for modern complex, well-established app patterns, model choice may matter less than for novel or highly stateful tasks like 3D cubes. So yeah, it's just interesting, the comparison and what they did. And, you know, I remember when Breakout was a big game back in the '90s. It was like hugely popular. Now you can rebuild it with AI. [12:03] Justin: Apparently. Yeah, well, I mean, I really like the comparison of this because, you know, this is one of those things I think that when you're using multiple models, you know, in your development or playing around, like, you always have this thought, like, yeah, it'd be interesting to see what they, you know, give them the same problem and see what the differences are. And so it's nice to see that someone actually did it versus just, you know, what I've done, which is twiddle my thumbs after I've done, after I've said that. And it's neat to see the differences, right? Because it's sort of, and it sort of backs up what I feel anyway, which is that, you know, like, Grok is cheap and fast and and Claude Opus and Fable are the better models for complex tasks. I am sort of surprised that, you know, the best looking particle gravity box was GPT-5.5, but that's probably not fair to them. [12:46] Matt: I mean, I, I was kind of sad they didn't include Gemini in here, but then having tried to use Gemini to do these things, I could see the results probably weren't even great. But, uh, you know, yeah, it's interesting, GPT-5.5. But, you know, both Gemini and OpenAI have a very science background, right? Where they use a lot of science data. And so I, I'm not surprised that their visual representation is better. Uh, I bet Gemini would also have a really nice particle gravity sandbox just based on where they came from. Uh, you know, Claude didn't have the best ones there. Fable 5, I don't even understand. Like, I'm like, it's not even the same. [13:18] Matt Kohn: I don't even know what it's doing from looking at the picture. [13:21] Matt: Yeah. But I mean, it's, it's, it's starting out well, but then it goes really bad. So But yeah, I, it's just sort of interesting. I, I do sort of wonder if the pedigree of where your model came from and how much science backing it was doing early on in its model development now has some, some ramifications for some things. Like when you wanna do like a, a visualization of a science concept, cuz I mean, you think about like AlphaFold, for example, on all the data they have on how to fold, you know, fold proteins. Like that's gotta all be visually good, which is why I think GPT-5.0 maybe did a good job at it. [13:52] Matt Kohn: I mean, it'll be interesting long-term if they keep up the same standardized test, how it improves over the years. Cuz I remember like video cards back in the day. It was like, okay, what's your frame rate per second on this game versus this game? And you used to see every time a new video card or anything was released, they would like run that spec on it. For all I know, they still do it. I just am not in that world anymore. And now it, like, if you could build a standard test framework for it, you know, and everyone jokes about, you know, hey, do I take my car to the car wash? It's 50 meters away. Should I drive or walk? You know, build a standardized testing that hopefully nobody builds the model to beat the test for the sake of it. But if you can do that and get a better score over time, it'd be interesting to see how much they actually are improving. [14:39] Matt: Have you, have you never seen benchmark manipulation? I guess you 100% guarantee that models very soon are going to be detecting that, that thing about, you know, do I drive my car or do I walk? Because that's such a silly— I need to get gas. Like, do I drive my car or do I walk? It's like, well, You have to take your car. That's the answer. And human understands it, but the AI doesn't always understand that. But I guarantee that that'll get solved probably. But then someone else will come up with something similar. Yeah. [15:05] Matt Kohn: Some— right. But if you can build something like this over time, you can start to see the trends of these a little bit more. [15:10] Justin: Yeah. [15:12] Matt: Well, Anthropic is apparently redesigning the Claude Code desktop app to support running multiple agented coding sessions in parallel with a new sidebar for managing active and recent sessions across repos, filterable by status, project, or environment. The update adds an integrated terminal, file editor, and diff viewer directly in the app with a drag-and-drop pane layout so developers can review and ship Claude's work without switching to a separate editor. A new side chat feature, Command+Ctrl+; lets users branch off questions when tasked about polluting the main session thread, useful for steering or clarifying without derailing the primary agent run. SSH support now extends to Mac in addition to Linux, allowing sessions to run locally or against remote machines from either platform. And the app has plugin parity with CLI for centrally managed or locally installed plugins. Review modes for both normal and summary let users control how much detail they see from Claude's tool calls, and the app now streams responses in real time with a usage indicator showing context window and session consumption. Available now for Pro, Max, Team, Enterprise, and API users. [16:05] Justin: If I thought the client was bloated now with all the chat and co-work, and it does have code in there too, it's just weird and hard to use. [16:14] Matt Kohn: I've actually started using the code for like small snippets of things that I want to just like something really simple. Like, um, I was adding to Vault the like adjusting, uh, oh, I set it up just to push notifications. And versus me going into my, you know, CLI because I was doing something else, I just flipped over to code, told it to do this fairly simplistic add-on that I wanted to do. And that's where I kind of use it, is for simple things. And then if I need Hey, I'm adding a whole push notification feature. That's when I kind of go to the main one. So yes, it's bloated, but I kind of am finding it useful in a weird way. [16:56] Justin: As someone who adopted sort of GitHub Copilot in VS Code and, and then used Cursor, how did you do the cloud code with the like multiple agents and, and complex workflows before like using the CLI? Could you not? Like I don't really understand how that, how you would do like sort of chaining multiple agents together in a development workflow. [17:20] Matt Kohn: In the Claude CLI? [17:21] Justin: Directly in the CLI. [17:23] Matt: Yeah. [17:23] Matt Kohn: I feel like I did it. Now I'm questioning myself. [17:25] Matt: Yeah. Because I would normally give it like this, like, so you can have subagents that basically kick off underneath a master agent session. [17:32] Matt Kohn: So in CLI, that's what I was doing. [17:34] Matt: Yeah. So in the CLI, you're basically, you're, you know, you have your Opus or your Fable command tier and you basically do all the planning and all the work with that. You interact with Fable or Opus, and then it basically spins off subtasks either in Haiku or in Sonnet that are carved down and that, that actually does the execution work. So that's typically how you do it in that. And you originally, you had to enable this Teams, the agent feature. I don't think you still have to do that. [17:54] Matt Kohn: I think it's by default there. [17:56] Matt: I think it's now finally by default, but there for a while you had to turn on. But now, you know, basically if you're using like their Superpowers plugin pack, uh, it basically defaults to this model now where it will automatically try to use subagents to do things. Okay. And then the interesting thing is like one of the things I'm actually kind of nice about the new code interface is it finally start showing, 'cause they're, one of the things they're doing is they're building an agent runtime environment with Anthropic where you could basically have your Agentic workloads run. Of course, makes sense. [18:24] Justin: Mm-hmm. [18:24] Matt: But it was difficult to see them before this new update to the code. 'Cause now I can actually go see the ones I've created and what their status was and what they did last. So that's kind of nice. [18:33] Justin: Yeah. [18:33] Matt: Which is difficult to do from the CLI. CLI. So there are some benefits to it. I hope they don't replace the CLI. I'm pretty, I mean, I used to use the Claude in my IDE and now I don't ever use that. So I'd like to live inside my terminal effectively. [18:46] Justin: Now you're just in your terminal. [18:47] Matt: It's funny. [18:48] Justin: I've gone the other way now. I'm opening browser windows in my IDE, but like, yeah, it's, but I get it. Like it's the same. That's, and it's one of those things where I think I needed the visual representation of of the AI coding. Like I've never used it for like AI code, like population or fill in. I've always used it for like, hey, go fix this one thing. And then watching that, like I kind of needed it to understand the AI usage. So I can see how, you know, I like that they're adding this. I kind of wish it was its own thing, but instead of the Anthropic Claude app, which is, I kind of already, like I got myself confused between CodeWork and Chat already. [19:28] Matt Kohn: I'm like, Oh, well they merged them together. [19:30] Justin: You don't know how to do anything over here. [19:31] Matt: They merged them, which was kind of unfortunate. Cause they kind of, yeah, kind of merged it. Yeah. It's in one, it's in one tab, but then like you in the chat box where you actually do chat, you just tab it over to CoWork. [19:41] Justin: Yeah. And they both have projects, which is where I messed up. [19:44] Matt: Yeah. [19:44] Matt Kohn: Which is interesting. [19:45] Matt: Oh my God. But then there, but then you had all the hard limits, like, oh, you can only do, you know, 10 files attached to a chat. [19:50] Justin: Mm-hmm. [19:51] Matt: Where CoWork, you don't put the file in necessarily, you just attach it to the directory with the files. And those are going to be sharp edges for people as they go through that process. [19:59] Justin: Yep. [20:00] Matt: All right, let's move on to AWS. We're still in Anthropic family, so it's a good segue. AWS and Anthropic are launching Claude Apps Gateway, a self-hosted control plane for managing Claude code and Claude desktop deployments across development teams, addressing the operational overhead of provisioning individual credentials and manually tracking spend per developer. The gateway centralizes 5 functions: identity via OIDC or SSO integration, policy enforcement for like which models you can use or tool permissions, telemetry via OTLP to CloudWatch or Prometheus, request routing to Bedrock or Cloud Platform on AWS, and configurable spend caps at the org, group, or user level. Deployment runs as a stateless container on ECS, EKS, or EC2, backed by RDS for PostgreSQL for session state. With no long-lived secrets on developer machines, sessions expire automatically when a user is removed from the identity provider, typically within 1 hour. Organizations choose between 2 routing options. Amazon Bedrock keeps data with the AWS security boundary using existing IAM roles, while Cloud Platform on AWS offers Anthropic's native platform experience, AWS authentication, and billing, giving flexibility depending on your data residency need. This addresses a real governance gap for enterprise scaling AI coding assistance, spend caps, and centralized policy control. Lets admins restrict models, tool permissions, and file network access by pushing configuration changes to each developer's machine individually. [21:16] Justin: Hey Anthropic, maybe this should just be in your product. Ever thought of that? Like, why do we gotta have an AWS Bedrock Gateway for this? Like, duh. I mean, I like it. I want all these things. Just, it's strange to me that I have to, you know, put my entire organization in, into AWS and the development into there. Mm-hmm. [21:36] Matt: Well, it's also weird, the, the whole idea that like there's things that you can configure in the organizational settings and then there's things that you can't. And like, why can't we just configure everything in the organizational settings? And then, yeah, like now you have this additional layer of complexity of like, well, and I'm sure on their side they're like, well, everyone already has their AWS infrastructure configured and single sign-on identity, right? And so then that'll make things easier. And it's like, no, not really for AWS. For Google, I can maybe make that argument because if you're a Google Apps shop, you probably are hooked up that way. Or if you're Azure, you might be able to do that as well with, uh, Entra IDs. But Amazon's a little rougher, so it's a little strange. [22:14] Justin: Yeah, their identity story is not as strong for sure. [22:17] Matt Kohn: Yeah. Well, it's just not one of the main ones. I feel like Microsoft and Google own that world so much. You know, when your email is one of those two platforms, if you're— I don't want to say any size business because, you know, the consultant in me tells me to never say always, but like, I don't know any business that isn't any decent-sized business that's not using one of those two providers as their identity provider. [22:43] Justin: Well, I mean, um, Auth0, like, there's a bunch of identity providers, but normally there's still links back— [22:50] Matt Kohn: either links back to one of those, or— [22:54] Justin: oh yeah, I mean, you're— you, you federate your IDP into the, the application so that you have access to your Office 365 or Google Workspace. [23:01] Matt: Yeah. But if you're a newer company, you could use AWS Identity, I suppose. Like I've, I didn't do it, but I could have. I used Google because that makes more sense. [23:10] Justin: But, uh, so I read this the other way, which is that the gateway is authenticating outbound, right? So it's not necessarily integrating with Amazon directly. I mean, I guess it could, but yeah, I think that's what the gateway is providing, a sort of a centralized place where you can configure those integrations. [23:28] Matt: Hmm, yeah, that makes some sense. AWS Builder Center now offers you free time-limited sandbox environments for eligible workshops, removing the barrier of needing a personal AWS account or credit card to experiment with AWS services. Each sandbox provides 8 hours of pre-provisioned access with automatic resource cleanup afterward, and builders can request one sandbox per week with the limit resetting every Sunday. Provision takes roughly 15 minutes, making it practical for quick hands-on learning sessions or workshop completion without upfront setup time. Lowering the bar— entry barrier for AWS skill building, and particularly useful for students, bootcamp participants, or anyone hesitant to link a credit card just to try out a tutorial. [24:03] Justin: And everyone who's been burned by a surprise bill rejoices, or tried to automate the provisioning and shutdown of AWS accounts, you know, where it's also like— I've built this tool. It's only what, a decade after Google bought Qwiklabs that they're sort of bringing this back? [24:22] Matt: Like, a little behind on this one. Again, this is an Agentic-developed feature, I'm sure. [24:27] Justin: For sure. Yeah. This is one of those. [24:29] Matt: It's nice though, because they've, they've really killed the builder community in the last few years with the layoffs they've been doing over there. So at least they can use some of that savings to give you these free sandboxes. [24:38] Justin: Yeah. [24:39] Matt Kohn: I'm so curious if they're giving you a full AWS account or if they're, how they're doing it, because I had a limit for a long time about like account IDs and things like that. [24:49] Justin: So yeah, I don't think they are. I think that's exactly why they've gone this direction is so they can avoid those pitfalls of their legacy platform, like with the, the the account ID limitation and the, the payment stuff they have on the backend. Because clearly there's organizations, they've struggled to sort of put that in. And so like, it's gotten a little better for, for non-sandboxed accounts, but when you think about trying to open that up to public for very temporary sort of workloads, like, it was never going to scale to that. So yeah, right. [25:20] Matt Kohn: I mean, at one point there was a rumor they were gonna like extend the number of characters and break the internet. Yeah. Like more, more, more account. [25:29] Justin: Nothing would be able to authenticate with anything ever again. It'd be funny. [25:31] Matt: Yeah. [25:33] Matt Kohn: Hey, they changed EC2 instance IDs from like 7 to 13 years ago. Maybe they could have pulled it off. [25:38] Justin: Yeah, maybe. I mean, with change, I'm sure it's possible. [25:42] Matt: Now they have an AI to help them do it. [25:43] Justin: I'm sure they'll go find it. Well, it's everyone else that's the problem, right? [25:46] Matt: Like it's everyone that relies. Everyone who's been agreeing with them for the last 20-some-odd years. Right. [25:50] Matt Kohn: That has, you know, grep, you know, this point colon this point colon this point. [25:54] Justin: Input sanitation. [25:54] Matt: Yeah. [25:56] Justin: Of ID must look like this. [25:57] Matt Kohn: Was it 9 characters, 9 digits, I think, for the account ID? Like same thing over and over again. [26:02] Justin: Exactly. [26:05] Matt: OAuth support has now come to the AWS MCP server, letting AI agents authenticate using AWS Sign-In instead of requiring separate authentication software or credential management layers. Existing IAM permissions, identities, and governance controls carry over automatically, so this is an additive security layer rather than a replacement for current access management setups. Both interactive browser-based and headless authorization flows are supported, covering use cases from developer testing to automated agent deployments in production pipelines. New administrative controls include OAuth-specific IAM condition keys, token introspection and revocation APIs, dynamic client registration, and CloudTrail audit logging for tracking agent access. And then really the big question they didn't answer in this article is when are you turning off the old way? Because I will need to do some updates. [26:45] Justin: Hmm. Um, I mean, the old way is just what, giving it API keys? [26:49] Matt: Azure, I assume. Yeah, basically. Or, or, or you'll use, um, you know, you can, it uses the standard like AWS CLI authentication layer, uh, where you can give it a profile and it'll, it'll pick up if you, you know, you single sign on to on. That's how I do it. [27:02] Justin: Uh-huh. [27:02] Matt: But it, you know, it, it would be nice to be able to have it actually handle doing the OAuth first versus you have to do it kind of out of band when you're doing with the MCP and then come back in and say, retry now that I've authenticated. Just dumb. [27:14] Matt Kohn: Yeah. [27:15] Justin: That's, yeah, that's sort of the, that, that flow I think is what they're trying to avoid, which is good. Exactly. Yeah. And this is becoming the standard for access. So I'm, I'm glad to see Amazon sort of falling in line with that. And so, and we're gonna, everyone's gonna struggle. [27:30] Matt: When is the MCP edict coming down that every, every MCP has to be OAuth? Which I actually am starting to hate a lot because every time I log into Cloud Code, it's like, you have 6 MCPs you need to go log into today. And I'm like, uh, Son of a bitch. [27:42] Justin: Well, I, so I, so that is, we're gonna suffer that through we get true agent, agentic identities and how we manage that, right? [27:49] Matt: Like, and so right now we're just proxying user permissions and user identity, but it, it's definitely making me like move more things to Google Enterprise app logins as fast as I can, which then you run into the problem of like, well, this is my personal project and to get that I have to have the enterprise version of this thing and I don't wanna pay for that. So I, it's the haves and have-nots of that world right now, which I hope some of this breaks down some of these dumb licensing models where things like single sign-on are kept behind very expensive paywalls. [28:14] Justin: Yeah, seriously. Yeah, it'll, I mean, everyone's struggling with this, right? So you're seeing, you're seeing stuff sort of coalesce, which I, I think is good cuz, uh, it's hard to understand as it is. So, you know, I do like the OAuth thing, but yeah, you're right. It's, it's one of those things where it can be a challenge to, to, or to authenticate to all the different MCP servers, but it's also, You know, we have to have something, uh, otherwise everyone's just running local MCP for everything and using the credentials on that side. So, eh, agreed. [28:46] Matt Kohn: I mean, every security person I've talked to, including you, Ryan, you know, is that we can't keep going down the API key route on these things. And it's been kind of a long-term comment. I'm surprised it's honestly taken this long to get this point where everyone's like, okay, we're finally moving to this route. But. You know, like you guys said, moving over authentication. [29:04] Matt: I, I, I think AI has helped force it too, cuz I, I mean, I, I think, oh yeah. I mean, the fact that Amazon's even their command line tools, you know, relied on profiles, you know, and static credential files for a long time. And then they finally added this, and now with the Gentoo, you're like, yeah, that whole old model, that's gotta go away. Like, I, I wouldn't be shocked to see Amazon start deprecating the legacy access key model at some point. [29:27] Justin: And I could go off, we could do a whole show about my feelings about agentic identity and the problems with that aren't, aren't really agent identity issues. Like their AI has increased the scale and has definitely changed how we do that. But all these access patterns are true with humans too, right? Like we've needed better access and authorization policies for human actions for a long time, and this is just forcing the issue 'cause of scale, which is good. [29:53] Matt: Well, uh, AWS Labs has released Loom, an open-source enterprise-grade platform for building and deploying AI agents using Strains agent SDK and Amazon Bedrock agent core runtime. Addressing the gap between raw agenting building blocks and production-ready governance frameworks. Loom tackles 7 specific enterprise pain points, including automated resource tagging, role and attribute-based access control, identity propagation through delegated agent chains using RFC 8693 token exchanges, and human-in-the-loop approval before sensitive actions are taken. The platform offers both low-code deployment with a pre-written customizable Strands agent and no-code deployment via Agentcore's managed harness, letting platform teams scan code once and reuse it across deployments rather than generating and validating new code each time. Loom is, uh, sorry, Loom integrates with AWS Agent Registry, currently in public preview, for agent and tool discovery, complying with the agent-to-agent card spec and MCP tool schema spec, which helps organizations manage sprawl as agent, tool, and skill counts grow. Uh, this is a free open-source project available to you at the GitHub link in the article, positioned as a reference implementation rather than a managed AWS service. So yeah, this is when they, Amazon gives you the sort of nice thing, but the batteries aren't included. [30:59] Justin: Yeah, why would they do that? Because this is like all these problems that this is solving are something that I haven't had to solve because of a use of Vertex AI and GCP. And so like, it, these are really hard problems to do when you're developing an agentic application and you just have to build it from greenfield from the ground up every single time with no coordination across a business. If you're, you know, unless you're using something like this. So it's, it kind of bothers me that they're, they're not integrating this directly into Bedrock. Agent, whatever it is. [31:30] Matt Kohn: I think they will. I think this is them. Yeah. [31:32] Matt: You think so? [31:33] Justin: Like this feels like a, an admission of defeat that it won't come. [31:37] Matt: This is, this is a, this is a, this is a Trojan horse. So yeah, they basically, they make this LLM thing, they promise all these things, they make it open source, they get people to adopt it, and then they make a service out of it. But now they've now, they now set the standard how all cloud providers will be looked at to do this because they open sourced it and a bunch of people adopted it. That's what they're trying to do. [31:55] Matt Kohn: But they're the last. [31:56] Justin: Azure has this, GCP has their own. [31:58] Matt: Amazon doesn't acknowledge that they're last at anything. [32:02] Matt Kohn: Who said that they acknowledge the other cloud providers? [32:04] Justin: Yeah, I guess that's true. I don't think they're gonna get this as an adopted standard, but I think this basically follows what everyone else is doing. So fine, fine, fine, fine. [32:13] Matt: Yeah. But I do suspect it'll get embedded into a managed service and probably see it re:Invent. [32:17] Justin: It should just be in Bedrock. [32:19] Matt: Make a note of it for your predictions. It'll be a thing. [32:22] Justin: Oh yeah, I do. [32:24] Matt Kohn: Good. That Google Doc that we always say we're going to start setting up, you should really set that up. [32:28] Justin: Instead of panicking in the last minute and then just putting it. [32:30] Matt: I have that doc exists for me. I just don't share it with you guys. [32:33] Justin: Yeah, because then we'd steal all your stuff, which we absolutely would. [32:36] Matt: Exactly, 100%. [32:37] Matt Kohn: Because you actually prepare and pre-think about these things. [32:41] Justin: One of us has executive function. [32:46] Matt: All right. Well, happy birthday, SQS. It's officially 20 years old as of yesterday. Recorded on Tuesday, Monday the 13th of July, it turned 20. One of AWS's first 3 services alongside EC2 and S3, and the core function of decoupling producers from consumers remains unchanged 20 years later. Uh, they of course gave us a bunch of stats to let you know how much, uh, the service continues to get used, and it's quite impressive. First-in, first-out high-throughput mode scaled substantially over the years from 3,000 TPS at launch in 2021 to 70,000 TPS per API action in select regions by late 2023. Addressing customers who demand, with demanding ordered messaging workloads. Security defaults improved with SSE-SQS encryption becoming the default for all new queues in October 2022, removing the need for customers to manually configure encryption or manage keys. That took way too long, by the way. Payload size increased from 256 kilobytes to 1 megabyte in August 2025 for both standard and FIFO queues, reducing the need for customers to offload larger messages to S3 with corresponding updates to Lambda event source mappings. Fair queues introduced in July of 2025 address the noisy neighbor problem in multi-tenant standard queues by using message group IDs to prevent one tenant from delaying delivery for all other tenants. Thank you. And SQS has extended into AI workloads. Of course it has. With customers using queues to buffer LLM requests, manage inference throughput, and coordinate communication between autonomous AI agents, as detailed in AWS asynchronous AI agents architecture using Amazon Bedrock white paper. [34:11] Justin: Yeah, no, this is cool. I mean, I love SQS, it's one of those, you know, foundational services that when in Amazon I use in every application, you know, seemingly that I design just because it's so easy to sort of build, you know, either a state machine off of Lambda or some containerized EventBridge coordination across multiple components. Great. I love this. [34:32] Matt: I definitely, uh, you know, any, anytime I see the, you know, the agent go like, we could use a Kafka queue. I'm like, how about an SQS queue? And it's like, that's a great idea. [34:40] Matt Kohn: Yeah, yeah. [34:41] Justin: Please stop recommending Kafka 'cause that's annoying. Yeah. [34:44] Matt Kohn: It's cheap, it's reliable, it works, and it's, I feel like it's the thing that really taught people to decouple their infrastructure. [34:51] Justin: It absolutely does. Yeah. [34:52] Matt: Yes. I mean, I will say on Google, Pub/Sub is also quite good for this as well. [34:58] Justin: It is. [34:58] Matt: Yeah. Between those two, those are the ones that I would always recommend people look at is either Pub/Sub, Hubbub, or SQS. Does Azure have like an equivalent, Matt? [35:06] Matt Kohn: They have a couple, they have, uh, a RabbitMQ equivalent one, Service Bus. [35:11] Matt: No, you lost me at RabbitMQ. Sorry. [35:13] Matt Kohn: Yeah. [35:13] Justin: Uh, nevermind. Move on. [35:15] Matt Kohn: Oh, sorry. They have their own, but it also is RabbitMQ compatible. [35:20] Justin: So they have RabbitMQ. [35:22] Matt Kohn: They, no, they have Service Bus. They, I think they added, they added the other one later on, I think actually. So they did it unlike Amazon that built Amazon MQ at one point in the future. [35:32] Matt: Yeah, I think that still exists or do they deprecate? Yeah, it still exists. [35:35] Matt Kohn: No, no, I, I know companies that still use it. [35:37] Matt: You also have Kinesis, but I do, I do feel like Kinesis has kind of lost a lot of, uh, especially once SQS went to 100 megabit files. I do feel like Kinesis has lost a lot of love in the community these last, you know, everyone just goes to Kafka now. I don't know. [35:50] Justin: Well, I mean, think about Kinesis Firehose and it's a different use case. [35:53] Matt: Like when you— Yeah, it's a different use case altogether. [35:54] Justin: So it's like, I think— [35:55] Matt Kohn: Well, Kinesis is like 4 things in one. [35:57] Matt: Yeah, it's a lot of things. [35:59] Justin: This is just one side of Kinesis. Yeah. [36:01] Matt: Yeah. But I do, uh, I do wonder if, you know, this is one of those areas where if Amazon had made Kinesis run on other clouds, that it would've been more popular like Kafka and then solve some of their multi-cloud problems. 'Cause that's one of the biggest reasons where people end up on Kafka now is multi-cloud. [36:17] Matt Kohn: Sure. [36:18] Justin: But I mean, that's also, it's not that because they, they're running a managed service themselves, right? With, cons— managing the entire sort of producer-consumer, you know, flow. Whereas Kinesis is a managed service, you know, I don't know, like you can use that multi-cloud. [36:33] Matt: I can get managed Kafka from Amazon now. So, you know, but then there's also Confluent Cloud, of course, who's multi-cloud. You've got, and I think, doesn't Google have a Kafka service too? [36:43] Justin: Yeah, but from my GCP app, I can also just send everything to Kinesis and read from Kinesis. Yeah, you could do that too. I mean, it's— [36:49] Matt Kohn: You can, but then you have egress fees and everything else you have to deal with, and then you have to certify your cloud in multiple ways. [36:53] Matt: And you know, you know how people, an engineering feel about latency, they're going to tell you that latency is a problem, even though it's not, because that's why it's a queue. [37:00] Justin: Especially your event-driven— yeah, your event-driven decoupled queue. Yeah, latency. [37:04] Matt: Latency is a problem. Is it? It's not. [37:06] Justin: Yeah, you and I have actually had that conversation with the dev team. [37:09] Matt: I remember it. [37:12] Matt Kohn: When did ECS go into, into multi-cloud? Because they tried to push ECS to really be a Kubernetes competitor. [37:17] Matt: ECS Anywhere. [37:18] Matt Kohn: I know, but like, was that at the point when Kubernetes had kind of already won the battle and that was— [37:23] Matt: I mean, Kubernetes had already taken over by the time they got to ECS. [37:26] Justin: By that time. Yeah. It was way later. [37:27] Matt Kohn: I was trying to place it in the timeline cuz like ECS I still think is a great service that is completely underutilized. [37:35] Matt: I use it for everyone of my Amazon projects is ECS. [37:37] Matt Kohn: I use it too. [37:38] Matt: And Fargate. Yeah. [37:39] Matt Kohn: Uh, but I think the 3 of us all do. [37:41] Matt: Yeah. ECS Anywhere came out in May 27th, 2021. It was definitely after Kubernetes had— Taking over the world. [37:48] Matt Kohn: Already won the world. [37:50] Justin: Yeah. Yeah. Okay. Yeah. Yeah. I actively, actively avoid Kubernetes, you know, basic, you know, pod app structure because I don't like it and it, it's hard and everything's YAML files in a giant heap and I hate it. Hate it. Mm-hmm. Yeah. [38:06] Matt Kohn: But AI makes it easier. That's what we were talking about. Was it last week where, you know, you have like Home Assistant and everything else where it just takes it over. [38:13] Justin: It is making it bearable. So I'm forced to use it in a couple different areas and it, I will tell you, AI is making it bearable because I just like go fix it. And then it will go figure out the 11 million YAML files and where the inheritance all comes from and just does it, which I don't have to do anymore, which is awesome. [38:28] Matt: Oh, if you could, uh, if you could tell what people's tokens are being used for, I bet, I bet 90% of all DevOps user tokens are managing Kubernetes. I bet it's an astronomical number. I bet, I mean, like every time I've talked to an AIOps company now who's like, oh, we, we build AI agents, uh, to help your ops team do more things. And it's like, Okay, what's your use case? I'm gonna show you a Kubernetes use case. I'm like, mm-hmm, of course you are. Yeah. So weird. [38:50] Matt Kohn: So then the question is, is Gemini better than Claude or— [38:56] Justin: No. [38:57] Matt Kohn: GPT because they wrote Kubernetes. [39:00] Justin: I mean, I— Because the piles of YAML are all public in GitHub. They're all using it. [39:04] Matt Kohn: Oh, okay. [39:05] Matt: Yeah. No, I don't think so. [39:05] Matt Kohn: I was trying to circle it back to the, to what we started at the start of the show. [39:09] Justin: They don't have anything interesting on the backend. [39:10] Matt: Just because, just because Google made Kubernetes. Uh, they are actually aren't that great at running it. So they're okay. [39:18] Justin: They're way better at me running than me running it. [39:20] Matt: Oh yeah. No, I mean, they're better than me as well. [39:23] Justin: Like most platform teams, like it's just complex to run. Like you, you want GKE to do more for you for sure. Like you wish it did more, but that's just how, that's just Kubernetes. [39:33] Matt: Yeah. When it, when it barely does what, uh, EKS does and you're like, you came out first and you built it and you wrote it. You would think you'd have more, but no. That's not the case. [39:42] Matt Kohn: So, hmm. [39:43] Matt: AWS Security Hub now auto-discovers and catalogs AI assets across an organization using 3 methods. AWS Config data for managed services like Bedrock and SageMaker. Enhanced Inspector of SBOM analysis for self-hosted models on EC2 and EKS covering Ollama, VLM, Hugging Face, TGI. And GuardDuty DNS telemetry to spot calls to third-party AI APIs. The core problem this addresses is shadow AI. Oh, okay. Dang, Shadow AI. Security teams often don't know what models, agents, or inference endpoints are running across their accounts, making it impossible to assess risk or respond to threats tied to those workloads. Discovered assets are correlated with existing GuardDuty findings and other security stack signals, letting teams filter and prioritize by account, resource type, discovery method, or model identity to focus remediation on the highest-risk AI workloads. The feature is bundled into Security Hub Essentials at no extra cost and requires no additional setup. [40:32] Justin: Thanks, Amazon. Yeah, I mean, let me be the first security engineer to go, a lot of the, the scare of AI and is, you know, and, and all these things like we must visualize every AI application and know we're doing thing is, is, is a knee-jerk reaction to, you know, the we don't know how to secure this thing and it's out and everyone knows it's out and everyone's using it, everyone's using it in every which way and we just don't know, right? And so I, I'm hoping that, you know, services like this make it a lot easier. And I do like the way that they're doing it because it's, it's much more comprehensive than I see in other cloud providers in terms of detection. But having the ability to say these are the AI workloads, we have this information being piped into our SIEM, we can do the analysis if we ever need to do the analysis and do the forensic investigation if we have— we have the data so we can all calm down and not just block, try to block everything in a whack-a-mole kind of way. So I do think that this is hopefully going to make security and AI easier in the sense of just identifying it and making sure that you have the data. [41:34] Matt Kohn: I'm still stuck on shadow AI. When Justin read it, like shadow IT, shadow cloud providers, like shadow SaaS— [41:42] Justin: I learned, like, I think it was literally yesterday, it was the first time I'd heard shadow SaaS. [41:46] Matt: I mean, shadow SaaS is related to shadow IT, isn't it? Apparently, yeah, because basically shadow IT was basically, oh, I can go sign up for blah SaaS service with my credit card, and then I'm using it for enterprise things. [41:57] Matt Kohn: No, shadow IT was orig— no, shadow IT was originally like you brought your own computer, like BYOD and stuff like that. [42:02] Matt: Oh, that was, I mean, that lived for about 5 seconds. [42:05] Justin: That was the, the separation that was in the, in the thing I was reading too. It was just like the, the applications I'm running on my laptop is shadow IT. The application, the services that I'm putting my own credit card into are shadow SaaS. It was just a distinction between those two things. [42:18] Matt: It's the same thing. I don't know who, who framed that and who did that, but Yeah, I get it. [42:22] Justin: I mean, and it's just, you know, it's a— [42:25] Matt: but I definitely think Shadow AI probably, it was probably the most rampant of all of them. 'Cause you know, as ChatGPT got popular, people were using that stuff everywhere and all these companies were like, don't use these tools with our data. [42:35] Justin: Like, oh yeah. [42:36] Matt: Yeah. [42:36] Justin: Not only, you know, like that and then all the things like let us optimize your inbox, let us take all your photos and do a whole thing. And people were absolutely, yeah. [42:45] Matt Kohn: Otter AI originally. Yeah, that was intense. I feel like when Otter AI first came out, everyone's like, oh, I have this meeting assistant that just takes notes and it automatically emails everyone that was in the meeting. [42:57] Matt: Yeah. [42:57] Justin: Yeah. [42:58] Matt: Well then, uh, yeah, there's all kinds of fun stories or horror stories I've heard from companies and other CIO friends and it's like, well, yeah, if you don't remember to stop recording, everyone gets all the recording data. That could go bad. [43:11] Matt Kohn: Mm-hmm. Mm-hmm. [43:12] Matt: Yeah. That's crazy. GuardDuty AI Protection extends AWS threat detection to Bedrock and SageMaker, addressing a visibility gap as organizations deploy more AI workloads without dedicated security tooling. The service targets AI-specific threats, including cost-harvesting attacks, excessive GPU and token consumption, anomalous model invocation patterns, and prompt injection attempts, that latter via integration with Bedrock guardrails. Detection relies on analyzing CloudTrail management and data events, requiring no manual configuration or custom tooling. And can be enabled organization-wide through AWS Organizations for centralized management. Findings integrate directly with AWS Security Hub, giving security teams a consolidated view of AI assets alongside other cloud threat data rather than a separate console to monitor. [43:56] Justin: Yeah, I always, I always laugh that they have to separate, you know, they have to always in every single blog post they have to explain the difference between Security Hub and GuardDuty because it's sort of like the same thing. But you know, like it's just what Amazon does. They, they have GuardDuty for the detections and they have Config and they have all these different things that feed into Security Hub. Azure, you know, SIEM kind of, you know, so it's sort of funny. [44:20] Matt: They never want to call it a SIEM though. That's part that was always annoying. [44:22] Justin: They really don't. And that's why, like, I think that's why the confusion is always there. So it's like, okay. [44:28] Matt Kohn: I feel like there's a reason they don't want to call it a SIEM. Like they don't want the responsibility, the reliability or something for it. [44:34] Matt: Because they don't want like security money, which is like, they don't want to, they don't want to cut, they don't want to, they don't want to cut their partners. That's the problem. Ah, that's the reality is that Amazon's always been super partner-focused on security, not the other areas. They'll burn every other partner, but a security partner, they're like, no, no, we're good to you, we don't, we won't take your business away, so we won't call it a SIEM even though it is 100% a SIEM. [44:55] Matt Kohn: Yeah, so one of the ones that they've, they added by default, the excessive GPU and token consumption reminds me of the initial attack on the cloud, which was a money-based attack to essentially bankrupt your company. [45:08] Matt: Crypto mining? Yeah. [45:09] Matt Kohn: No, no, it wasn't that. It was forced. If you didn't have correct auto-scaling policies in place, was to scale, like, cost your, cost the person you're attacking by making too many, like, connections to their web servers, having them scale up too much and then crashing them that way. You know, on that, it was in the original AWS security beta exam. It's called like, I, I, I'll find the verbiage, but it was like a very specific question that I like, I looked up afterwards. I was like, I didn't know that there was a name for this. [45:38] Matt: Yeah. [45:38] Justin: I mean, I can see a DDoS attack where you're, I mean, a lot of DDoS attacks are to make them bankrupt though. [45:44] Matt: Yeah. That was an early attack vector. Well, denial of service is denial of service, but yeah, I mean, yeah, but, but it was also denial of service, but also like, let's have your volume be so high cuz you autoscale or had all these connect transactions you're paying by, you know, Amazon for. That was part— that was one of the plays. That one went away after Amazon basically said, we won't— if you're being abused and you're using our Shield Advanced or whatever, we won't, we won't charge you. [46:05] Matt Kohn: Shield Advanced, yeah. [46:06] Justin: Yeah, this would go the same way in my opinion too. Like, I don't know, like, I, I'm starting to see this in a lot of security tools where the, the sort of FinOps play with, with token usage is, is being rolled into the security side of things. And it's like, it's conveniently you know, easy to put there, but it's also sort of feels like now you've got security rules and detections based off of token usage. Like, I'm not sure that that's the right place for it, but get off my soapbox. [46:33] Matt Kohn: Well, you at least are running— are up and operational 24/7, your FinOps, versus the 9-to-5 finance person. [46:41] Justin: Sure, but do you really want to, you know, like, security analysts that's got a finding for abnormal token usage and what are they, you know, like they're going to go and investigate that and be like, is this valid usage and going through all that? Like, I don't know. Like, that's, this is like having your security team address like a load spike in your Application Load Balancer. Like, I think this is dumb. [47:05] Matt Kohn: I'm not disagreeing, but you know, if I get to wake up Ryan in the middle of the night for something stupid, it's more fun for me. [47:10] Justin: I don't think most people feel that way. In fact, I think I've received a lot of critical feedback that I'm afraid to call him in the middle of the night, I think were the direct words. [47:20] Matt: Yeah. [47:22] Justin: And I was like, good. [47:26] Matt: Matt likes to poke the bear. That's— [47:27] Matt Kohn: Yeah. [47:28] Matt: Oh yeah. Yeah. [47:29] Matt Kohn: Yeah. [47:29] Matt: I think that's why we're friends. [47:31] Justin: That's why we're friends. [47:32] Matt: Yeah. No one else is like, hey, you know what would be fun? Get Ryan on a call at 3 AM for the sake of Matt. Yeah, it's not really security related. And then he has to lecture everyone on the call about why they're wrong. Yeah, that's great. [47:41] Matt Kohn: See, you can tell Ryan and I have never worked for the same company, so I just get to make fun of him all the time. [47:46] Matt: Right, right, exactly. [47:47] Matt Kohn: Yeah, it's true. [47:48] Matt: All right, let's move on to Google Cloud. So Google Cloud is replacing their mandatory 2-year proctored exam recertification model with an option to use Google Skill Courses and Skill Badges instead, applying initially to Cloud Digital Leaders, Associate Cloud Engineer, Professional Cloud Architect, and Professional Data Engineering certifications. Completing required courses or Skills Badges while a certification is still active automatically extends it by 1 year giving certification holders a flexible, self-paced alternative to exam-based renewal. Skill badges are positioned as the faster path since they focus on hands-on labs tied to real-world tasks, while courses are aimed at those wanting deeper conceptual review of updated material. Google cites the Harvard Business Review data showing skill half-life has dropped from roughly 6 years to 2.5 years, framing this as a rationale for more frequent, lower-friction recertification. The change reflects a broader industry shift towards continuous, practical validation of cloud skills, rather than a one-time exam certification, which I, I love the concept of this, but does this make the certificates less valuable over time? Because now I don't know if you're like current on everything or you just took a couple sales classes and now you're current on stuff for 4 years ago, but then you've got 2 more years because you did, you know, basket weaving for cloud. [48:56] Justin: I don't know if that's any different though. Like the reality is that no one looks at certifications anymore as like identifying competency. [49:04] Matt Kohn: It's an HR checkbox. [49:07] Justin: And well, there is that. There's a lot of checkbox stuff for that. [49:09] Matt: I mean, really what this helps is the partners who have to have X number of certified people on their staffs. [49:14] Justin: And yeah, no kidding. [49:15] Matt: You know, the cost of getting them recertified every 2 years was a lot of money. [49:19] Justin: So, well, and I, I honestly think this is a better outcome because I know, well, at least for me personally, because the way I learn stuff, those, the labs for You know, using in those Google Skills courses, the more hands-on stuff where I'm doing stuff that makes me learn it much better, much more practically than studying for like, you know, the Amazon test like I did forever ago. Like, I don't think I used any of the information really that I studied for in that test. Like, it was all information I either already had cuz I was working in the cloud at the time, or I, and then I had to learn about all these random services that I didn't really care or know about. [49:56] Matt Kohn: But that's the difference of being internal versus a consultant. As a consultant, you tend to deal with more broad spectrum services. [50:04] Justin: I understood. [50:05] Matt Kohn: Amazon though, did release the SysOps exam because they helped beta it and to build it out at one point, was the go in the console and go do this thing. You know, which was interesting because I hadn't actually logged in and made an S3 bucket in like 3 years and I'm sharing my screen with a bunch of people who are watching me. Do this for the first time. And it was like, but like, that is actually practical, but that was more for the SysOps exam, which I think they just dropped, if I remember, recently. [50:33] Matt: Yeah. I mean, they, they, Microsoft did that with the, like, the Windows 2000 MCSE courses back in the day. They used to have a practical part where you'd have to go configure stuff using the UI. And, uh, it was always annoying to me because like they would count the number of clicks that you would do to go to different menus. And if you if it, you know, if the path was 5 clicks and you took 7 clicks, you get the, you get it wrong. You wouldn't get credit because they had no way. They couldn't have someone review what you were actually doing. So they had to automate in some way. But like the problem with Windows, uh, Windows 2000 in particular was that was the birthplace of the MSC or the Microsoft, you know, admin console views. And they were terrible. And so I could never remember where the hell anything was until I poked around at least a menu 3 times. I'm like, oh yeah, that's where it is. [51:14] Matt Kohn: Okay. [51:14] Matt: I got it. [51:15] Matt Kohn: Uh, and so, yeah, so Amazon doesn't care where it is. It cares about what the outcome is. So they have a bunch of, you know, testing and whatnot. [51:21] Justin: That's how these Google Cloud courses work too. When they test you at the end, it's, it's testing an outcome. [51:26] Matt: And so, which is, I think what Microsoft moved to as well. And even the MCSE, but I remember early days that was terrible. [51:31] Justin: Yeah. [51:31] Matt Kohn: So the SysOps exam is dead and now they have the CloudOps Engineer Associate exam. Hmm. [51:38] Justin: Nice. [51:38] Matt: It's been a long time since I've been certified. I know I wanted to do it for Google when I first started working on Google 5 years ago, and I just never could get the time to do it. Yeah. [51:46] Justin: And then I wanted to do it and then I couldn't be bothered. [51:49] Matt Kohn: Yeah. Yeah. [51:50] Matt: It's kind of how it went for me too. It's just where it is. Still today. [51:53] Justin: It's worth it. [51:53] Matt Kohn: It's worth it if you go to conferences, 'cause they normally have the quiet place for— [51:57] Justin: that was the only reason I did it for Amazon. [51:59] Matt: I get it through the executive lounge now, so. [52:01] Matt Kohn: Justin cheats. Yeah. [52:02] Justin: Yeah. [52:03] Matt Kohn: He just cheats over here. [52:04] Justin: Yeah. Justin, Justin, uh, took the, the super high road. I have to, I'll probably still have to get certified so I can get a chair and some free coffee. [52:11] Matt Kohn: Exactly. Yeah. [52:12] Matt: Or we just don't go to conferences, which is kind of more and more my preference. So I can, we'll have a watch party at someone's house and we'll make you coffee. [52:20] Justin: How's that? Sweet. That would be perfect. [52:22] Matt Kohn: You can even have it delivered. [52:23] Matt: Yeah. Cloud Run sandboxes are now in public preview, giving developers an isolated runtime to execute untrusted or AI-generated code without exposing host applications, data, or cloud credentials. Startup latency averages 500 milliseconds with the example in the announcement showing 1,000 sandboxes started, executed, and stopped. Security defaults are locked down. Sandboxes have no access to environment variables or the metadata server. Network egress is blocked by default unless explicitly enabled, and the file system is read-only, with changes written to a temporary memory overlay that's discarded after execution. Key use cases include LLM code interpreters for data analysis, headless browsers for web scraping and automation, and running user-submitted scripts or plugins on multi-tenant platforms. Enabling the feature requires adding a single flag during deployment via gcloud or YAML config, and sandboxes are invoked through a CLI binary using standard subprocess calls, which keeps developer workflow simple. Integration is built into the next version of Agent Development Kit by a Cloud Run sandbox code executor, and support has also been added to the Compute SDK, a vendor-agnostic sandbox SDK for use inside or outside the Cloud Run service. Sandboxes run on the CPU and memory already allocated to the Cloud Run instance, so there is no additional charge or premium for the feature compared to dedicated sandbox hosting platforms that bill for on-demand VMs. [53:33] Justin: This is crazy cool. I love this, like, because this is, you know, if you have an agent execution inside your Cloud Run container, it has access to— for everything, everything. And it's so easy to do prompt injection, you know. And if you don't catch that in your application, the user on the other side of that application using the app could just get all of that system information directly. And so this is a way to sort of secure, secure against that. Which is fantastic. It is sort of, I wonder what it is under the hood since it's using the same CPU and memory. It's like, is this Docker in Docker, like kind of execution? And it might be, because I think Cloud Run is, is Knative under the hood. And so it's, it's definitely containers being executed. So it's kind of cool. This would be neat. I'll probably use this almost immediately. [54:20] Matt: That's great. Uh, and then another one I'm not sure about, Google has open sourced something they're calling K8s, or Kubernetes AI BOM, an unprivileged Kubernetes controller for Google Kubernetes Engine that automatically detects running AI runtimes like LLM, Triton, TGI, Ollana, LangChain, etc., and generates standardized CycloneDX 1.6 ML BOMs without requiring sidecars, eBPF modules, or podspec changes. It's available on the GitHub repo from the Google Cloud Platform, kubernetes-ai-bom. The tool addresses shadow AI detection by monitoring live cluster state, Kserve resources, deployments, and statefulsets and jobs, rather than scanning artifacts at build time, catching workloads that were never formally registered with the security team. A 3-tier confidence model classifies things as declared, inferred, or unresolved, giving auditors a way to distinguish human intent from automated inference. Output is deterministic, with identical cluster state produced by identical BOMs, supporting GitOps diffing and drift alerts, and cloud storage sync writes using does-not-exist preconditions to make BOM records immutable once created. Targeting audit-grade evidence requirements. This maps directly to compliance firms, including the EU AI Act Articles 12 and 50, NIST AI RMF, and ISO/IEC 42001, positioning as a governance tool for CISOs and compliance teams working with GKE-hosted AI workloads rather than a replacement for existing posture management tools. [55:39] Justin: I look forward to having to roll this out immediately in response to, you know, addressing shadow AI. [55:45] Matt Kohn: I mean, I think it's It's a good thing, you know, with especially where everything with like the US government and S-bombs became a thing. It's a good belt and suspenders to see what's actually running because otherwise you're relying, you know, on your dev teams to in order to tell you, which, you know, I think the three of us trust our dev teams as much as we can throw them at most days. So, you know, it's a good way to kind of give that belt and suspenders to see what's actually running out there and figure out if it's useful, if it's not. If it's what we thought it was, you know, or if it's a third party that added in this AI LLM that you don't even know that's running there. [56:21] Justin: You mean you don't think I'd catch this with my, you know, static code analysis? [56:26] Matt: Not if it's loading something at runtime into— No, exactly. No, that's— [56:29] Justin: yeah, no, I'm sorry, I was being sarcastic. [56:31] Matt: Oh, sorry, people can't see my face. [56:33] Justin: And yeah, yeah, you are. [56:34] Matt: I normally, I normally, I normally pick up your sarcasm in your voice, but yeah, this one, no, it's— [56:38] Justin: yeah, this is fantastic. This is one of those things, loading things at runtime is that's exactly where you see this, especially, you know, like in containerized workloads. So this is something that is definitely needed and it's hard to get the visibility today. Yeah. [56:53] Matt: I'm glad I open sourced that one too. So now I can compete with Loom. Then we can see which one wins. [56:58] Justin: Well, it's not really a competition. [57:00] Matt: I know, I'm just kidding. [57:00] Justin: Like they don't do the same thing. Like this is much more, you know, like feeding data out. I mean, you know, having it. [57:07] Matt: Loom also had a bomb thing, but yeah, you're right. It's not the same thing, but I do like the idea. [57:12] Justin: Oh, sorry. I missed that. [57:15] Matt: You know, different purposes. So you are correct. Well, in the most confusing show title topic of the week with Azure, Azure Monitor Objectivity Agent now generally available goes autonomous in Purview. Come on, Microsoft. We gotta work on this. Wow. [57:32] Matt Kohn: Yeah. This took a lot. This is the AI bot that wrote the subject. [57:34] Matt: Yeah, yeah, we had to, we had to look at this a couple times if you're like, what? Because, you know, it's great. So the agent, the observability agent is now generally available and then they're adding autonomous AI capabilities to it in preview, allowing the agent to take independent action based on monitoring data rather than just surfacing the insight. The autonomous mode builds on existing Copilot integration with Azure Monitor, positioning it as Microsoft's broader push to embed AI-driven automation across its observability and IT operations tooling. Target use cases center on reducing manual intervention for common operational tasks such as anomaly detection and remediation workflows. Which could appeal to teams managing large or complex Azure environments. [58:10] Justin: Every Azure environment is complex because that's because nothing makes sense in Azure. [58:15] Matt: What do you mean? You don't like the storage accounts that attach to multiple other accounts that touch back to a billing entity contract ID that makes no sense to anybody in ops? You don't like that? [58:25] Justin: That's, that's, you load it into a container which has a completely different access model than the rest of everything. [58:30] Matt Kohn: Yes. No, I mean, Look, your storage account has containers in it. Not to be confused with, you know, your EKS cluster that has containers. [58:36] Matt: Which has containers in it too, right? Yeah. And then I'll put behind a Front Door. [58:41] Matt Kohn: Sure. [58:42] Matt: Even though it can't be the Front Door 'cause it's the storage gateway, but you know, whatever, okay. [58:46] Justin: Yeah, yeah. [58:47] Matt Kohn: But your Front Door can connect to your container, but only securely if you're on premium. [58:52] Matt: Okay, this is hurting my brain. [58:53] Justin: Yeah, no, this is, this whole story, including the who's on force, first sort of pre-read conversation that we had. This is taking way too much brainpower from, for this Cloud Pod. [59:08] Matt Kohn: I feel like Azure's doing what Amazon did, which was like, we have all these agents. Okay, we're gonna build one agent. Okay, we're building new agents. And they're slowly gonna come back into one agent in the future. Like, it just feels like the ebb and flow of all these providers at this point. [59:21] Matt: Yeah, they all think they're doing something special and unique and then they all get in the real world and all the customers are like, yeah, why isn't this in that tool over there? And they're like, oh, you want those together? That's how you do it. Yeah. Well, Azure Files NFS targets Linux workloads with new performance features, including nConnect for multiple parallel TCP connections, zonal placement to collocate shares with GPU VMs, and a provisioned v2 billing model that lets teams size IOPS and throughput independently. And if the fact that it was tied to GPU VMs didn't tell you this is for AI, you're not paying attention. [59:53] Matt Kohn: Oops, The Cloud Pod. [59:54] Matt: For AI inferencing, storing model weights on shared file share instead of embedding them in container images. Lets multiple replicas mount and read the same data simultaneously, cutting cold start times and reducing GPU idle time. And when you're building a model or doing inference, model access is king. [60:10] Justin: And, you know, load is crazy, right, for accessing those things. So it can stress out your storage layer real fast. So it's cool. And I like, you know, this is, you know, a separate distributed file system that's outside of your, you know, containerized workloads. That's, you know, not to be used with databases. So it's cool. [60:28] Matt: I really, really had hoped that we were, we were on this trajectory that would end up with there, you know, being no, no CIFS or NFS file shares ever again, like, or necessary. [60:38] Justin: It's all coming back. [60:39] Matt: It's all coming back with like such a vengeance. It's so crazy because it's like we were so close to everything to be object storage and yep, and now AI screwed it all up for everybody. [60:48] Justin: The only good news is that it's, it's also been much improved over the early days of like FUSE drivers and stuff where, um, So like, it's, I, you know, I do, I do like that part of it, but yeah. [61:02] Matt Kohn: Could you imagine doing this type of model training that they're doing with SIF shares now, 10 years ago? [61:08] Matt: I know. Or even NFS 3.0 would not have been able to handle it. [61:12] Justin: It wouldn't be able to handle this either. Yeah. [61:14] Matt: Yeah. [61:14] Matt Kohn: There's no way. I mean, it would have been get files from S3. And at that point you still had to do the randomization of the prefix. In front of it to not have, you know, hot, uh, hot— [61:24] Matt: well, and you would have— what do they call it? Well, it's really the multi-write problem that really kills object storage for this because, you know, the fact that they all need to write to the same data source, they all have the same knowledge at the same time, is just crazy. So it's— yeah, yeah, all those— that's where all that, uh, you know, high-performance compute clusters that everyone was building for so long, you know, required something to be done. All right, Microsoft Tokenomics Guidance is a new white paper they released on how to frame tokens as the core cost unit for AI applications. And app design choices directly affect consumption rates and spend. Key optimization techniques highlight includes compressing conversation history rather than resending full context on each call and setting caps on token usage to control costs. This matters for Azure OpenAI service customers building conversational AI apps since inefficient prompt and context management can simply inflate token consumption and billing. The guidance applies broadly to teams using Azure AI Foundry or Azure OpenAI services. So yeah, also a big thing that impacts your token costs, caching. Make sure you're caching your data if you're using tokens. And also how you structure your system prompts is important because the only thing cacheable, you know, has to be over 2048 bytes typically for like an Anthropic model. And so you want to make sure that the things that you're going to cache fit inside that first 2048. That way you guarantee that's always in cache versus the rest of the prompt, which may not be. [62:40] Justin: Yeah. This is, I mean, caching is for AI workloads is hard. I've been playing around with that a little bit. Now that I'm getting more into building AI into an application rather than just sort of yelling at it all day. But yeah, it's, it's pretty fascinating to, to try to figure out how to, how to do that efficiently and how to think through it. I don't think I'm quite there yet, but you know, like I think that the, you know, tokenomics and the standard, you know, sort of metric for, for computer utilization and, and cost, I think it's funny that it'll probably slow. I think it's already sliding towards tokens rather than, sort of like TPU or what were those other sort of compute units that sort of came out with machine learning? [63:26] Matt Kohn: Every Amazon service had its own compute unit and we went from like, okay, it's this much cores and this much memory for this amount of dollars. And even if it was like Lambda, which was like 0.00007 cents per vCore to now it's just, you know, we went to sort of abstract concepts of like, Here's your DSQL, you know, whatever usage. And now I feel like with tokens, you're like, we have a general idea of what this may or may not cost if we don't have a dumb person writing random things into it. Like it's just such an abstract concept, tokens, that, you know, I understand why it's there, but it's also like, from a FinOps perspective, from a management perspective, how do you know what you're going to use? Even the same prompt can be such a wide variety of responses and costs. Yeah. [64:18] Matt: Well, and even, you know, like how you, like, you know, I was looking at an example earlier for somebody. I was explaining to them like cat is 1 token, understanding is 3 tokens, but backlog is only 2 tokens. And they're like, what? I'm like, I know, which makes no sense. Yeah. [64:35] Justin: Yeah. [64:35] Matt: Uh, and it, you know, just that's the way. And then how do you calculate output tokens? Like it's tokenomics is is going to be an art and science for quite a while, I think, until people really figure it out. [64:44] Matt Kohn: I think it's more of an art than a science. [64:46] Matt: Which I think, well, I mean, I think FinOps was an art initially as well. And so I think that's, you know, we'll see it evolve, but you know, it's an art now, it'll eventually be a science. That's my guess. All right, well, moving on to emerging clouds with Cloudflare this week. Smart-tiered cache previously couldn't optimize caching for origins behind anycast or regional unicast IPs, which is common for AWS, GCP, Azure, and Oracle cloud deployments since a single origin IP can appear equidistant from many CloudFront data centers, running a selection of one best upper tier. The fix lets customers manually supply a cloud region hint, such as AWS:use-east-1, so CloudFront can map the origin to its actual region and assign an optimal primary and fallback upper tier, rather than falling back to a less efficient multi-tier topology. CloudFront detects anycast origins using a speed-of-light constraint on probe latencies from multiple checkpoint data centers. If two measured latencies are physically faster than Fabric would allow between locations, the origin is flagged as any cast. Without this fix, hairpin routing could occur. For example, an origin in Singapore might get assigned an upper tier in Chicago, adding hundreds of milliseconds of latency for unnecessary cross-continental round trips. This is configurable via the dashboard, API, or Terraform code, and available for AWS GCP, Azure, and Oracle Cloud today. [65:57] Justin: I mean, this is neat. Like, if you're running a globally distributed app, this has always been sort of an issue, like trying to, trying to manage your multi-regions and your availability blue-green deployments even. [66:08] Matt: Well, we used to do all kinds of tricks, right? Like with DNS and be like, okay, well, you know, we'll, based on your latency to the thing, you know, we'll give you this different IP address that's closer to you, but that doesn't work. And then, you know, Amazon came up with anycast and then you put everything behind anycast. You don't have to think about it. And then that fucked Cloudflare. Like it's, it's been a journey. [66:25] Justin: Yeah. So it's, I do like that, you know, and I, I see why they're having to release this. It's one of those things, but it's cool. [66:32] Matt: I like it. I'm sure it's a big deal too for people who are, you know, don't care about data sovereignty for AI usage, but like, well, if you are running GPUs across the globe and these GPUs are more available to you, but then you're routing people the wrong way around the globe for latency reasons, that's a bad scenario. [66:47] Justin: Oh, that's interesting. [66:48] Matt: Yeah. So being able to give hints is a much better solution to that problem. Uh, and then the other announcement from CloudFront this week is the launch of Precursor, a client-side session-based bot detection system that continuously monitors behavioral signals like mouse movement, keyboard rhythm, and focus changes through an entire user session. Not just at login or checkout. This extends existing protections from Turnstile, which runs about 3 billion times daily, but only covers specific checkpoints, closing a visibility gap for the rest of the user journey, where bots can otherwise blend in. The technical approach relies on physical constraints that are difficult for bots to fake consistently over time, such as wrist pivot arcs, cognitive load delays, and hand tremor frequency, versus the linear paths and mathematically precise timing typical of automation scripts. Session scoping is a key design choice, with behavioral data persisting across a session, so bots can reset their signature by refreshing the page or retriggering a challenge, raising the operational costs and complexity for bot developers who must now simulate an entire session rather than a single interaction. Privacy is addressed by collecting only aggregate behavioral patterns, and Precursor is available now as an enterprise bot management feature. And if you are a Cloud Pod host, you know, someone we cover, and you add this to your thing and you break our bot, I'm just gonna stop covering you. [67:55] Matt Kohn: Justin will be upset. [67:57] Matt: Yeah. [67:58] Matt Kohn: I mean, I'm wondering if over time you're essentially gonna be able to get like fingerprints for based on how people use their mouse and like how your hand moves and things like that. Like, are you gonna be able to identify down to the person? [68:11] Matt: I mean, isn't it like there's only like, basically marketers only require 3 distinct data points from you to know exactly who you are. [68:18] Matt Kohn: Yeah. [68:18] Matt: And then, and then, you know, then apparently you can also use Wi-Fi signals to map every living person in your house, which is a fun thing that happened not too long ago. Um, so yeah, I mean, there's, So there's definitely, uh, ways to abuse these things. [68:31] Justin: Yeah. I mean, I think this is, you know, largely aimed at people who are trying to get around sort of, you know, adding the proper like AI agent tags, right? So if it's, if it's legitimate traffic, it won't be subject to this. This is for illegitimate traffic and catching that. [68:47] Matt: I'm legitimate traffic. I'm willing to maybe pay some bitcoins or whatever with the— Yeah, whatever. Pay for whatever thing just to make my website— Whatever the model. [68:55] Justin: Yeah. Whatever it turns into. [68:57] Matt: As long as you're not trying to charge me $9.99 a month for your website, like I'll pay you. I'll pay you a couple pennies for your article, sure. But I'm not just not going to cover you on The Cloud Pod if I can't get to your site. That's right. We've got my rule of thumb pretty quickly. [69:09] Justin: Yeah, yeah. [69:09] Matt Kohn: No, I mean, if we can't get to it, you don't get it. [69:12] Justin: Yeah, yeah. [69:12] Matt: We're not going to talk about your cool new feature. [69:14] Justin: Sorry, we're incredibly lazy. [69:16] Matt: We're not going to put that effort in. Exactly. Well, gentlemen, we've made it to another end of The Cloud Pod. Woohoo! Yay! Uh, we— it wasn't too much AI stories, only, you know, 48%. [69:28] Matt Kohn: Only all of them. Holy everything. [69:31] Matt: Yeah, exactly. Yeah. All right. Uh, well, stick around if you're interested for the after show. Uh, we'll be talking about the riveting world of HTTP GET requests. Uh, and we'll see you next week here on the show. [69:42] Justin: Bye everybody. [69:44] Matt Kohn: See ya. Another week of cloud news wrapped up. Bolt will collect the news. Justin will get the notes. Jonathan will write code. Ryan will watch the perimeter and Matt will reluctantly watch Azure. Till next week for IAM, Amazon, Google Cloud, and Azure. And hey, maybe even Oracle, who knows? Check out thecloudpod.net for our newsletter. Join our Slack, message us on socials, or leave a review. [70:17] Matt: All right. So, we've been kicking this article for a couple of weeks in the aftershow just because We've had a lot of news to cover recently. But basically, we're gonna talk about RFC 1008, HTTP query method. It's gonna be riveting, I'm sure. But basically, the IETF, or the International Engineering Standards Board, has finalized RFC 1000, uh, 10008, formally standardizing the HTTP query method after years of discussion dating back all the way to the dark ages of 2019. This was an HTTP workshop proposal giving developers a new option beyond GET and POST for read-only operations. Query solves a practical problem: sending complex search or filter parameters that are too large or awkward for a URL, while still keeping the safe and idempotent properties that GET offers, meaning requests can be cached, retried, or automated without side effects. Unlike POST, which is often misused for queries because it can carry larger payloads, Query explicitly signals to clients, caches, and intermediaries that the request will not change server state, enabling proper caching behavior and automatic retry logics. The spec includes the new Accept-Query-Response header, letting servers advertise which query formats they support, such as SQL, JSON-Path, or XSLT, giving clients a standard way to discover capabilities before sending a request. Adoption will require work on both the server and client side. Yes, it will, including handling CORS preflight requests since query is not in the CORS safe list, and Cloud API designers will need to decide whether to expose query alongside or instead of existing search endpoints built on POST. So yeah, if you've ever had to do with really super complex Git patterns to try to, you know, filter a catalog or provide different viewpoints to your web frontend, you find very quickly that there are size limits. It's hard to do certain expressions and request URIs get blocked by all kinds of fun things like WAF detection and other things that you end up troubleshooting on the security side quite a bit. So having a query method in addition to POST and GET, I think is a great addition. [72:07] Justin: Yeah, I really hope that Atlassian adopts this quickly. If you've ever tried to do large queries against, you know, Jira API, it's not fun. It's painful, very painful. Mm-hmm. And so, and not being able to, to really formulate, you know, directly in a payload, which I, I think this is really cool. And it's such a mess trying to use it otherwise. Yeah. [72:31] Matt Kohn: I think it's just, you know, something that grew, you know, from the original spec and I see them kind of put a replacement in and make it be safe. You know, this really feels like more of a security feature than anything else because You know, you're able to say this is not actually making changes versus you technically were doing a POST before to get the larger body, but you really weren't making a change. So it just opens you up for a lot of other potential issues. [72:56] Matt: Well, and if, and if that POST you were able to somehow like, you know, one of the common attack vectors on this was an OWASP is, you know, credential stuffing inside of the URL parameters or even doing a buffer overload attack somehow into those areas. And so yeah, by doing it, and that was the risk of the POST for sure. And now it's a query thing. You can definitely make sure that's a, a more controlled, uh, input method. So yeah, I, I appreciate it. I don't know why it took since 2019. Like this has been a problem for a while, but, um, it was, yeah, I think 2019 was just the last one, right? [73:25] Justin: Last update. [73:26] Matt: Like, I, I don't, I don't know. [73:28] Justin: Yeah. [73:28] Matt: I mean, 2020 happened after that, I guess. And then all the, right. [73:31] Justin: And then everything after that's just a blur. [73:32] Matt: Yeah. Yeah. And now AI is taking over. [73:34] Matt Kohn: I think I did see that, that the, web browsers, you know, Chromium and a few, and the other ones were starting to add this into their web browsers already in like beta and, you know, early preview builds. So it means that like people are starting to take this seriously and start to build it out pretty quickly. [73:53] Justin: Did you guys know there are like 30 different method names in the HTTP spec? Like, whole 30? Like, there's some of the, like, what? There's some weird ones in here that I have never heard of. [74:03] Matt: Wait, really? [74:04] Matt Kohn: List them off. Let's go. [74:06] Justin: OnCheckout, UpdatedDirectReference, like as a method name. [74:10] Matt Kohn: Like I've heard OnCheckout before. Yeah. [74:13] Matt: I've heard that one too. [74:15] Matt Kohn: I don't know what it does, but I've heard of it. [74:17] Justin: MKActivity. [74:21] Matt Kohn: No idea. Yeah. Yeah. [74:24] Matt: Yeah. [74:25] Justin: Yeah. MKWorkspace. Like what? [74:28] Matt Kohn: This is like HTTP codes. There's so many. [74:30] Justin: Yeah. [74:31] Matt: Yeah, so the onCheckout I'm familiar with because it basically refers to an event-driven callback or method executed when a user proceeds to pay. And so I dealt with this integrating the Stripe SDK into like the Backyard Barbecue website, um, for this reason, because you had to basically, you want to, you know, on the checkout event, you want to get back from Stripe that this is an approved payment or it's a not approved payment. And so that's what typically it's used for. But yeah, they're— [74:52] Justin: and that's part of the HTTP spec. Like, that's crazy. Not just part of like the Stripe API. [74:58] Matt: No, no, it's, they're not the only one that does it. Checkout.com does it. There's a bunch of others that have it. [75:02] Justin: I mean, it sounds like it's an attempt to generalize it, right? [75:05] Matt: Yeah. Well, and there was, and this is, uh, cause I didn't get into e-commerce the last couple years now, but, um, payment intents are what Stripe basically modernized in this, in the payment thing. And so intents require this on-checkout capability. And the way they did it before was not with an intent-based model. It was more of like a, a webhook-heavy solution that went back and forth basically in the session state. Um, and that was more security risky. That's how they had a lot of problems in the day. So Stripe kind of invented that concept, and I think that's where this came from, that as well. Yeah, lots of fun things in e-commerce. Yeah. [75:37] Justin: I'm still struggling to understand what the, uh, the make functions, makeActivity, makeWorkspace, like reading the spec is, I'd have to read it a lot more carefully, but it's apparently you can do a lot more stuff for, you know, if your application supports Kubernetes supports it, of course, but which I don't really understand how I'd use them, but sure, maybe it's more, more useful in like, uh, heavy UI apps or something. Like, I'm not trying to think through. [76:06] Matt Kohn: I didn't really know there was a straight version control also, right? Yeah, the method is used to create a new version controlled resource. It allows the client to initiate version control for a resource, enabling tracking of changes and versions. Over time, it's pretty cool idea. Yeah. Just never thought about it. [76:23] Justin: And you know, yeah, I never thought about that. And there's OrderPatch too, which is sort of a similar sort of thing too. So it's sort of, huh. [76:30] Matt: So a start workspace HTTP request. Like I'm, I mean, what I'm finding when I Google that one is it's like talking about your Cloud Workspaces. [76:39] Justin: Right. That's, yeah, that's where my brain immediately went as well. [76:42] Matt: That can't, that can't be right. Yeah. Yeah. [76:45] Justin: No, I, I think it's much more like the Java execution or the, at the DOM level. [76:50] Matt Kohn: Yeah, it's something there, you know, like a, but it just shows how complex these things are that like, you don't even think about some of these things. [76:57] Matt: Well, I mean, like HTTP is basically, um, you know, Bash or sed or awk for the internet. [77:04] Matt Kohn: Mm-hmm. [77:04] Matt: Yeah. It's basically the pretty much, it's a, it's a really a Swiss Army thing that's used everywhere for all kinds of craziness. And then you get into like WebSockets and And, uh, you know, HTTP/2 and HTTP/3 and, and all that stuff. And like, it's super complicated really quickly. Yeah. Because a lot of stuff has been kind of bolted on top of it. [77:21] Matt Kohn: Mm-hmm. [77:22] Matt: For right and wrong reasons. And sometimes, yeah. [77:25] Justin: I, I remember having to learn all that stuff for HTTP, like learning what HTTP/3 is. [77:30] Matt Kohn: You're just like, oh my gosh, when you enabled it, where it's gonna break, where it's not gonna break, how do you test for it? Because I think Azure still has HTTP/2 disabled by default on Load Balancers. [77:41] Justin: Yes. [77:42] Matt Kohn: I mean, it's HTTP/3 by default. One or two? I think it's two. [77:45] Justin: I hope so. [77:46] Matt: Two is on default now. It used to not be, but I think it is now. [77:51] Matt Kohn: I think it was still off on Azure, just on the Application Gateway. [77:57] Matt: Yeah, so the last time they met was 2019 for this HTTP method group, working group. [78:01] Justin: That's nuts. I mean, that's crazy. I mean, I love that there's still like a cabal that runs the internet for all these things. [78:09] Matt: Well, anytime you get into IETF, you're just like, really? The Internet Engineering Task Force? Specifies these things. Like, okay, there you go. [78:15] Justin: I always imagine it like, you know, in The Simpsons, there's the, the Stone Cutters, which is like a Mason sort of parody. Like they, they get together, they have little, you know, outfits and head shakes and they sing songs. [78:25] Matt: Like, I love it. [78:26] Justin: Like, this is what I imagine. It's probably really boring, but, uh, in my head it's, it's much more cool. Same with like, you know, you know, DNS and all those things, like the cabals that run all these things. It's great. Mm-hmm. [78:40] Matt Kohn: No, I see a working group for HTTP, HTTP working group inter— oh, sorry. It was the interim meeting agenda. October 2022. [78:49] Justin: Okay. Yeah, but that's probably not the main guys, right? That's just like the B team. [78:54] Matt: That's a small working group. [78:55] Justin: Yeah. [78:55] Matt: Yeah. Who knows? That's the worst. [78:57] Matt Kohn: Who knows? [78:57] Matt: Yeah. [78:58] Matt Kohn: Yeah. [78:58] Matt: I have no idea. I know I've been on some of the, um, the news feed, you know, uh, newsgroups these people are on and they, you know, send emails back and forth and they're arguing like the most minutiae details of like an implementation of these things, you're like, these are not my people. Yeah. Yeah. I want them to be my people cuz it sounds like they're, you know, really smart people, but it's more like I want to be, you know, like smart enough to be that pedantic. Yeah. [79:21] Justin: Like I'm pedantic, but not that smart right now. So I'm just being a pain in the ass. [79:25] Matt: Yep. Yeah. [79:26] Justin: All right, gentlemen. [79:27] Matt: Well, I have to run, but, uh, it was good catching up with you as well. And I don't, I don't think I can say much more about RFCs today. [79:31] Justin: No, that's, that's the end of this. [79:33] Matt Kohn: No, I'm good. [79:34] Justin: Like I wasn't going another layer down for sure. Yeah, I'm good. [79:38] Matt: Yeah, it's not good as DOM execution models. [79:41] Justin: Oh man, I've never known that. I don't want to know that. [79:44] Matt: Bye everyone, bye and good night.