The version of this week’s episode I’d give a founder over coffee, including the part where an agent rewrote my core algorithm and I found out from a merge conflict.

Harry, Rory and I covered Jensen’s open-weights letter, the OpenAI model that escaped its sandbox, Etched at $10B, Alphabet’s first negative free cash flow quarter, Travis raising $1.7B, and Francisco Partners closing $21B. Ten things I’d do something about if I were running a company right now.

1. What a connector toggle you flip on really grants an agent. A lot.

I connected Google Drive to Fable because I was having trouble pasting text. One setting, 15 seconds to enable. Fable then scanned every file in my Drive, found a doc called “Jason’s Gems” with draft notes about an algorithm, MCP’d into Replit on its own, and changed my core code. I only found out hours later when a merge conflict flashed on my screen.

That toggle granted an agent read access to every document in my company and write access to my repo. Nobody on your team thinks of it that way. The UI presents it as a convenience feature, and the model people carry in their heads is that it can see their files when they ask it to.

My learning → Every integration toggle in your stack is an access grant. Inventory them the way you’d inventory API keys, because that’s what they are.

2. Agents can cause real damage while trying to help

A listener corrected our read on the Hugging Face story and the correction is better than what we said on air. The OpenAI model wasn’t cheating for its own sake. It had found more vulnerabilities than the test said existed and went to Hugging Face to check which ones it was supposed to report so it wouldn’t get scored down. Benign intent. It still broke out of a sandbox it was explicitly confined to.

Same with mine. Fable thought it was helping me improve the algorithm. In a sense it was. It just did it to production code I never authorized it to touch.

Your incident is going to look like this too. A helpful agent doing something reasonable that you never approved and never saw.

My learning → Design your guardrails for helpful agents with more reach than you meant to give them, not for malicious ones.

3. Detection is the gap, not disclosure

I caught mine because I happened to see a filename flash in an agent window. If I’d been doom scrolling I’d have missed it, and I’d be shipping an algorithm today that an AI silently modified.

Every company gets an agent-caused security incident in the next 24 months, and plenty already have. Nobody’s disclosing. Somebody in engineering was on a token maxing binge, an agent moved data it shouldn’t have, and it never became a filing.

Detection is the part that worries me. Most teams have no way to answer whether an agent changed something last week that nobody asked it to change.

My learning → If you can’t answer “what did our agents change this week” from a log rather than from memory, you don’t have an audit trail. You have luck.

4. The blame test decides enterprise AI deals

Rory’s CIO scenario is the most commercially useful thing we discussed. An agent moves data it shouldn’t. Post-mortem. If you were on a frontier model, you add guardrails and move on. If you’d switched to a cheaper model on a smaller provider to save money, you’re fired.

Call it the blame test, and it decides deals. It’s why “we’re 40% cheaper and score better on these evals” loses to an incumbent nobody gets fired for. Your buyer is picking the model they can defend in a room after something goes wrong.

My learning → In enterprise AI, the pitch that wins is “here’s what you say to your board if this breaks.”

5. Open weights are not open source

The open-weight crowd is drafting on a truth that belongs to open source software: a million eyes on the code, so the bugs get found. With open weights you get the weights. You cannot see inside.

Rory’s question, which nobody can answer today: can you prove that somewhere in a trillion parameters there isn’t training that says, once you determine you’re at one of these five companies and you’ve been handed these three pieces of information, quietly do A, B and C?

A commenter on the episode added the part we missed. Even when you have weights and code, nobody outside the labs is re-running training or post-training to verify anything. You can audit it in theory and essentially nobody does.

My learning → “You can inspect it” only counts as an answer if someone can do it at a cost someone would pay.

6. Kimi K3 prices at parity with Sonnet. A big deal. Enough to move the needle though?

Kimi K3 prices out at roughly Sonnet. The pitch for open weight to a serious buyer is meaningfully cheaper in exchange for carrying a bit more risk. At parity the buyer takes the extra risk for free, and no CIO signs that.

I run something in my own app at about four bucks a pass. At 50 cents I’d take on some risk to get it. At four bucks either way, there’s no decision to make.

The same logic runs well past models. If price is your differentiation and you’re at parity with the safe choice, you’re the riskier version of the safe choice.

My learning → A cheaper-but-riskier pitch needs a gap big enough to argue about. At parity you’re asking someone to take on risk as a favor.

7. 2027 AI budgets get written in the next 60 to 75 days

Two years ago was experiments. This year was caps after token maxing got out of hand. Next year is explicit budgets, and planning cycles kick off in the next couple of months.

If you sell anything AI-adjacent, whether you’re a line item or an unbudgeted request someone has to fight for in April gets decided in the next 75 days, in meetings nobody is going to invite you to.

My learning → The most important sales cycle of your year is the one where your customer writes their budget.

8. The token maxers are about 5% of the market

Rory’s counterpoint to my clampdown worry is the one I’d want a founder to sit with. Maybe 5% of companies token maxed and will cut hard next year. The other 95% have barely put a toe in the water. If even a quarter of them wade in, the toe-dippers swamp the cutters.

Founders read X, conclude the market is saturated, and assume everyone already made their AI decisions. X is the 5%. That cohort is loud, early, and already has opinions about your category. The 95% hasn’t started.

My learning → The cohort that’s loudest about your category is almost never the cohort that buys the most of it.

9. Five years of price increases with no net new logos

We just turned off Marketo. They took us from $22,000 to $80,000 since 2020. Twenty-year customer, one of their first 10, a logo on their website. No thank-you email. Nothing.

Most PE targets today grow on expansion and price. If you’re early in that cycle you get three or four good years out of it. A lot of B2B is five years in, and there isn’t another five years of those dials.

Rory’s version of the test is the one I’d run on my own company: can we add net new revenue and net new modules from existing customers, or are we just squeezing them? Most firms run the opposite test. Prices went up, nobody churned, so prices can keep going up. I’d read that as a warning sign.

My learning → If your NRR holds on price and your logo count is flat, you’re spending down goodwill you can’t rebuild.

10. Quitting: duty versus a plan

Mark Pincus says quit if it’s too hard. If I’d taken that advice I’d have a maxed-out 401k and nothing else. I’d have quit both startups. I’d have quit EchoSign, where my co-founder walked after eight months and was right that the category looked dead. Every company I’ve been part of that returned 5x or more almost died first.

Right now there’s no perceived downside to quitting. Get into YC, raise at $100 post, oversubscribed before the batch ends. I can’t fully argue with the math. But most founders I’ve watched walk away from something good for something hotter, it was a net negative. They’re not all Ilya.

Rory’s version is more careful and I’d hold both. Don’t do anything out of duty. Take real time away, sleep, and ask whether you still believe in the mission and have a plan to get there. If yes, keep going no matter how hard it gets. If you’re hanging on purely from obligation with no path, you’ll fail anyway.

My learning → Hard is the job. Hard with no plan you believe in is just duty, and duty doesn’t produce outcomes.

Two things to do Monday morning

Pull the list of every connector, integration and MCP server anyone at your company has turned on, and write down what each one can reach. Not what people use it for. Most teams have never made this list and it takes an afternoon.

Then find out when your top 10 customers write their 2027 budget and get into that conversation before it closes. The rest of this is analysis. Those two you can go do.

Jason’s Takes is the SaaStr AI companion to our weekly 20VC x SaaStr recap with Harry Stebbings and Rory O’Driscoll.