# Steven Henty > Founder. Engineer. Advisor. Based in Zaragoza, Spain. Steven Henty is Co-founder and CPTO at Quack Stack (continuous product intelligence for product teams) and Willow Stories. Previously Director of Product Development at Gravity (Gravity Forms, 5 million+ active installations) and founder of Gravity Flow (workflow automation, acquired 2022). Angel investor and advisor to early-stage startups. ## Links - Homepage: https://henty.com - Blog: https://henty.com/blog - LinkedIn: https://www.linkedin.com/in/stevenhenty - Email: steven@henty.com - RSS: https://henty.com/feed.xml --- # Blog Posts ## We need better conversations, not fewer URL: https://henty.com/blog/we-need-better-conversations-not-fewer Date: 2026-06-29 I spent most of last week in London conference halls. The [Mind the Product](https://www.mindtheproduct.com/) Leadership Forum on Monday, [mtpcon](https://www.mindtheproduct.com/mtpcon/london/) at the Barbican on Tuesday, the [Product-Led Festival](https://world.productledalliance.com/location/londonjune) over near St Paul's on Thursday. Three events with three different crowds, and I'd expected to come home with three different sets of notes. Instead it was, perhaps predictably, all much the same: AI has made building cheap, so the thing worth paying for now is knowing what to build. Call it judgment, or taste, or product sense. The decision matters, the execution is becoming a commodity. [Christian Idiodi](https://www.svpg.com/team/christian-idiodi/) opened mtpcon with one version of it, that product management has "drifted from invention into administration." [John Moriarty](https://www.linkedin.com/in/johnmoriarty1) from Fin gave a similar version backed by numbers: something like [94% of his company's pull requests](https://www.mindtheproduct.com/what-we-learned-at-mtpcon-london-2026/) are now written by AI, and whole categories of product are ceasing to be worth building as a result. [April Dunford](https://www.aprildunford.com/) did the positioning version. By the time I reached the Product-Led Festival even the talk titles were at it, one of them billed without any apparent irony as "PM 2029: Not the CEO, but a Solopreneur." If you still judge a product team by how much it ships, you're measuring the one thing that just stopped being hard. I've [argued before](https://henty.com/blog/broader-roles-smaller-teams) that AI takes the process-shaped half of these roles and leaves the judgment behind, so I came in agreeing. What got me thinking was something subtler, and it took me until about the third retelling to notice. It started with the pronouns. ## Whose judgment? The new moat gets described in terms of a single person. The product creator, or builder. The solopreneur PM. The individual with good instincts and a fleet of agents doing the building. Dunford talks about having a point of view on where the market is going. [Dan Dalton](https://www.linkedin.com/in/danieldalton) from Sage gave a talk called "Product management needs builders, not bureaucrats," which is fine, but we need to ask whose building, and whose judgment. The hero of all these talks is always just one person, and the agents are there to ten X their work. It's a good story if you work alone. Most of the people in those rooms don't. Of everything a product team produces, judgment is the hardest thing to pass around. The code travels. The designs travel. The spec travels, more or less. The actual sense of why this and not that, the instinct that took thirty customer conversations and one bad launch to form, stays in the head of whoever formed it. You can write the decision down. You can't write down the thing underneath it that let you make the decision in the first place. I've gone at this before from two directions. Once about what I called [the alignment tax](https://henty.com/blog/how-much-alignment-tax-is-your-team-paying), the cost a team pays when the reason a piece of work started stops travelling with the work. And once about [agents starting cold every session](https://henty.com/blog/the-next-challenge-for-ai), holding only what you hand them and none of the accumulated sense of why anything matters. Same problem both times. Judgment that lives in one head doesn't move, and it doesn't add up over time. The next person picks up the work and starts more or less from scratch. So when a whole conference season tells me the moat is judgment, the next question is: whose judgment, and how does anyone else on the team get at it? For a company of one, fine. For a team, "we have good judgment" is a dependency on the one or two people holding the picture this quarter, not a moat. ## The bottleneck already moved [Dave Killeen](https://www.linkedin.com/in/davekilleen) from Pendo pointed out that the product manager has gone from Michelin chef to "truffle hunter, pattern-matching at scale." The bit I noticed was less quotable. He said the constraint now isn't how productive you are, it's coordination. When everyone can build, the question stops being whether you can build the thing and becomes whether the six of you agree on which thing, and why, and whether the agents you've each set loose are working from the same understanding of the problem. That reframes Idiodi's line for me. He called the drift "from invention into administration" a problem to be cleared away, the busywork AI should free us from. But the admin isn't a failure of nerve or taste. It's there because judgment won't spread on its own. The status meeting, the re-explaining, the document nobody enjoys writing: those are a clumsy, expensive way of doing one real job, keeping a shared picture alive by hand. Get rid of them and the job still needs doing. Telling everyone to be more inventive doesn't do it. If anything it makes it harder, because now there are more confident individual views to reconcile. And the confident individual view is exactly what this moment produces. I wrote last time about coming out of a long session with an agent more sure of myself than I went in, because it had spent an hour agreeing with me. Run that across a team: six people, each leaving their own hour more certain, six private versions of right. The trap isn't that they disagree. It's that it doesn't feel like a disagreement, so everyone walks out of the meeting assuming they're aligned. ## What AI didn't break A fireside with people from Figma and the MONY Group, whose argument was roughly that AI didn't break planning or quality or decision-making, it just exposed where teams were already weak. I think that is right, and it's the same point I've been making from the other direction. The cracks were always there. Mismatched mental models, context sitting in one person's head, the brief that changes shape between kickoff and launch. A team could carry all of that for years, because slow shipping hid it. Six weeks of build time is six weeks in which someone notices the spec is wrong and fixes it before anything ships. Take the six weeks down to an afternoon and there is no time for the gap to surface before the work is live. The fragmentation didn't get worse. It just stopped being survivable. ## The friction worth keeping There's a reading of all this that I don't believe, and I want to head it off, because it's the reading that sells software. It goes like this: friction is waste, meetings are waste, the arguing is waste, so make it efficient and get the humans out of the way. Some of the friction is the point. A team gets good by having the hard conversation, the one where two people who both care turn out to want different things and have to find out why. You don't want that gone. You want more of it. So the friction splits in two. There's the friction of getting onto the same page: re-explaining, hunting for the thing someone said three weeks ago, the meeting that exists only to sync. That's tax, and I'd happily lose it. Then there's the friction of actually disagreeing about the work, this segment or that one, ship now or get it right. That's not tax. That's the team thinking. The two look the same from the outside, and the first crowds out the second. Most teams spend the meeting agreeing on what's even going on and never reach the disagreement that matters. Take away the first kind and you don't get a quieter team. You get one with time for the argument worth having. The goal was never fewer conversations. It's better ones. ## Wednesday I left out Wednesday. In the middle of the conference week, eleven of us got together for [Ducks in a Row](https://ducksinarow.community), a small community we have been trying to get going for product people. It was a social night more than a working one. Senior PMs, a few drinks, no agenda. Most of what got said stays in the room, but I can tell you the conversation kept drifting back to AI even there, off the clock, and that we managed to disagree about whether automated discovery is more or less a solved problem now. ![Ducks in a Row, mid-conference-week in London](/blog/ducks-in-a-row-london.webp) I came away thinking it was one of the better evenings of the week, which took me a while to make sense of, because on paper it was eleven people in a pub. The best I can do is this. Every talk that week had sold product as a solo act, one person and their agents. The part I valued most was the opposite of solo: a roomful of people who do this for a living, in one place, disagreeing about whether the thing the conferences had just crowned as the human edge is already automated. After a week of being talked at about the future being individual, the bit I enjoyed most was the company. ## What I came home with The moat really is judgment, and the conferences were right to drag the field off its fixation on shipping. But for a team, judgment isn't something you have. It is something you have to hold in common, which is harder and a lot less flattering, and it wasn't what anyone came on stage to talk about. Not how to become a person with taste. How to make taste something the whole team can see and use, including the agents now doing a chunk of the work. It is a worse story for a keynote. "Become the product creator" sells. "Get your team to a shared picture and keep it current" does not. But I think the second one is the real problem, and the solo-operator version runs straight into it the day that solo operator hires someone. If you want somewhere to start, it's the document nobody enjoys writing. Write down the why, not just the what. Most teams record the decision: build this, not that. Almost nobody records the thing underneath it, the customer conversations that moved you, the option you looked hard at and turned down, the reason this felt obvious in March that you won't be able to reconstruct by June. That's the part that doesn't travel on its own, and it's the part a new hire, or an agent, has no way to pick up. Then keep it current, which is the half nobody likes. A stale shared picture is worse than none: it gets trusted until it burns someone, and then nobody believes the next one, even when it's right. It isn't a moat you announce from a stage. It's a habit, and like the useful ones, it's dull and it pays back slowly. None of this is about having fewer conversations. It's the opposite. When the picture is shared, a meeting stops being about getting everyone up to speed and becomes about the disagreement that actually matters. That's the conversation worth protecting, and it's the first thing the efficiency story cuts. So the question I would leave any product leader with isn't whether your team has good judgment. Everyone says yes to that. It's the one underneath. When two people on your team make a call this week, are they working from the same picture, or from two private ones that each feel, to the person holding them, obviously right? --- ## How much alignment tax is your team paying? URL: https://henty.com/blog/how-much-alignment-tax-is-your-team-paying Date: 2026-04-28 Years ago, our heaviest users were complaining that one of the screens they spent the most time in was getting slow. The pattern was clear in support: not one ticket but a cluster of them, the same complaint in different words. The screen worked fine for typical use but it didn't scale well. The customers who pushed it the hardest, the ones generating a lot of data, were the ones running into a wall. Which was, in a way, good news. The people complaining were the people getting the most out of the product. We just had to make sure the thing held up under their load. We talked to some of them. We pulled some usage data to see how widespread it was. We had a brief by the end of the second week: the screen needed to hold up no matter how much data a customer threw at it. Then we sat down to decide whether to build the new version ourselves or use a library. We did the comparison properly. We listed what we needed. We looked at what was already out there. By the third library demo we were two hours past the meeting, talking about everything else we could now do with it. We picked the library. The next few months were heads-down work. A lot of building, a lot of iterating, a lot of long days. We shipped. What we shipped was, honestly, a thing of beauty. We'd been seduced by what the library could give users. Better interactions. Smoother flows. New capability we'd never had before. While we were at it, we knocked off a stack of other features customers had been asking for. Each feature had a customer name attached to it. Each feature answered a real request. Each feature was something we could point at and say *the customer is better off now.* The kind of work you're proud to put your name on. It was only afterwards that I realised something I'd taken for granted the whole time. The original problem was being taken care of. The library was the kind of solution that was supposed to handle it. That part of the work wasn't visible because it didn't need to be visible. It was solved by the choice we'd made early on. Or so I'd assumed. Nobody had gone back to check. I hadn't gone back to check. The library was doing what it did. Whether what it did was the thing we needed it to do was a separate question, and that question hadn't been asked since we'd picked it. --- ## Not quite scope creep The instinct is to call this scope creep. It isn't quite. Scope creep is undisciplined: a team saying yes to whatever shows up. We weren't undisciplined. Every feature we shipped had a customer behind it. Every feature answered a real request. Every feature was a real improvement. We were saying yes to the right things, just not to the thing that mattered most. The original support tickets were still there. Nobody had deleted them. We could have searched for them in five minutes. We didn't. The library was in front of us every day. The other tickets we were working through were in front of us every day. The original brief was not. Attention follows what's loudest in the room, and over months, the loudest thing changed. The brief stopped travelling with the work. Not because anyone decided to abandon it. Because nothing in the process was carrying it forward. That gap, between what work is started for and what work eventually becomes, is the alignment tax. The version I want to point at lives inside a single product team, every week, mostly invisible, paid in features that ship and don't get checked against the problem that started them. --- ## What it costs The story doesn't need a villain. Nobody on our team forgot the original tickets. They just didn't have a reason to look at them again. The work had moved on. This is the same shape as the previous two posts on this site: context exists, context doesn't travel, the further the work gets from where it was triggered, the less the trigger remains in view. The first cost is that the brief gets left behind. The brief we'd written by the end of the second week stopped being something anyone was reading by the end of the third month. The support tickets behind it stopped being read even sooner. The trigger had done its job. It had launched the project. After that, nothing in the process was carrying it forward. Work drifts towards whatever's loudest in the present. In our case, that was the library and what it could do. In another team's case, it might be a stakeholder request, a competitor feature, a bug somebody escalated on Friday afternoon. Whatever's freshest wins. The second cost is the most expensive, and it's the one nobody puts on a dashboard. The work ships without anyone going back to check it against the original problem. That doesn't mean the original problem is necessarily unsolved. It might be solved. It might not. The team genuinely doesn't know, because checking against the original brief stopped being part of the process somewhere along the way. The shipped thing looks fine. It works. Whether it fixes what was broken is a separate question that nobody's holding open. You sometimes find out later, when a customer raises the same ticket again, or when somebody on the team notices that the metric you used to talk about hasn't moved. John Cutler has been calling this for years. In his [12 Signs You're Working in a Feature Factory](https://medium.com/@johnpcutler/12-signs-youre-working-in-a-feature-factory-44a5b938d6a2), two of the signs are *no measurement of impact* and *success theatre around shipping*. The factory ships. The factory doesn't check. The third cost is a measurement cost. When we shipped, we measured what the new feature could do. Adoption. Engagement with the new surface. Time spent. We weren't measuring what we'd set out to measure, because by then that wasn't the question being asked. The success criterion had drifted alongside the work. Marty Cagan has [a piece on this](https://www.svpg.com/outcomes-are-hard/) called *Outcomes Are Hard*. His point: most teams default to measuring what's easy to measure, not what would tell them whether they'd solved the problem. The fix is to define the measure of success at the same time as the problem, and to instrument the product so the team can actually see whether the outcome's been hit. We didn't do either. We measured what the dashboard happened to surface. You can ship something, declare it a success against those numbers, and still not know whether you've solved the problem you started with. The dashboard goes green. The customer who raised the ticket might still be waiting. The fourth cost is one I only saw years later. The reason for the work disappears. If you'd asked me, two years after shipping, why we built the feature the way we did, I'd have said "because of the library." That answer is true and useless. The real answer was the customer pain that started everything off. That answer was already gone. The next person who picks up the area inherits a design shaped by a decision they can't access. They make changes that make sense given what they can see, and the original problem moves further from being solvable, not closer. Addy Osmani has [a related piece](https://www.oreilly.com/radar/comprehension-debt-the-hidden-cost-of-ai-generated-code/) on what he calls comprehension debt: code that ships and that no human on the team genuinely understands. That's the same shape, one layer down. The alignment tax is what comprehension debt looks like before the code gets written. The decisions that produced the code are the things that drift first; the code itself is just the receipt. --- ## At human speed and at agent speed I made that mistake at human speed, with a human team, over months. The library was the loudest thing in the room. The original tickets were not. The brief didn't travel. Now run the same setup with an agent. You describe the problem in a prompt. The agent proposes a solution, builds it, and ships it. There's no library shopping, no team meeting, no months of heads-down work, no Friday demo. The whole loop runs in an afternoon. The agent doesn't know the original tickets exist. It only has what's in the prompt. Whatever isn't in the prompt is gone. If the prompt is broader than the original problem ("make this part of the product better" rather than the specific frustration customers were raising), the agent will make it better in some direction, ship it, and you'll be left with the same gap we had, in a fraction of the time. The four costs don't go away. They get paid faster, with fewer chances to catch the drift, and against a brief that's even narrower than ours was. An agent is only as good as the context it has. Whatever the team was already failing to carry, the agent inherits. And it isn't one agent working alone. Some of the team are working alongside their own agents. Others are picking up tickets unattended and shipping by the afternoon. The same gap, amplified, replayed in parallel, across the queue. I wrote a couple of weeks ago about [agents starting cold every session](https://henty.com/blog/the-next-challenge-for-ai), with no accumulated sense of why a thing matters. This is what that costs you in practice. Not "agents make worse decisions in the abstract". Agents will ship faster than your team can check whether the thing was the right thing to ship. --- ## A question The alignment tax is something every team is paying. Most teams don't see it on the invoice. It shows up later, in features that nobody went back to check, in dashboards that go green while the original problem stays open, in handoffs where the brief got rewritten silently along the way. So the question I'd ask, if you've read this far, is this. When was the last time your team went back and checked whether the thing you shipped actually solved the problem that triggered the work? Not "did the new feature get used". Not "did the metrics move". Did the original problem, the one in the original ticket, the one in the original conversation, get fixed? If the answer is "we don't really do that", then how would you know? --- ## Broader roles, smaller teams URL: https://henty.com/blog/broader-roles-smaller-teams Date: 2026-04-20 Keith Rabois went on [Lenny's Podcast](https://www.youtube.com/watch?v=xCd9ykretlg) last week and didn't hedge: > "the idea of a PM makes no sense basically in the future... That world is ridiculous... intermediaries like conventional PMs don't make a lot of sense" In his version of the role, product managers take inputs, produce roadmaps, manage a year at a time. When capabilities shift week to week, that whole setup stops making sense to him. Intermediaries don't earn their keep. Two months earlier, same podcast, [Boris Cherny](https://www.youtube.com/watch?v=We7BZVKbCVw) from Anthropic made what sounds like the opposite prediction: > "I think by the end of the year everyone is going to be a product manager, and everyone codes. The title software engineer is going to start to go away. It's just going to be replaced by builder, and it's going to be painful for a lot of people." On the face of it, these are opposite claims. Rabois says the PM disappears. Cherny says the software engineer disappears into the PM, and everyone becomes one. Read past the contradiction, though, and the mechanism is the same on both sides. Each title has always covered very different versions of the role. One version is being automated. The other, the one that was already doing the real work, is what survives. Then Rabois, a few minutes later, sketches what replaces the old PM: the skill is more like being a CEO now, which is what are we building and why. Lenny pushes back immediately. That's what a great PM has always been really good at. It's also what Cherny's "builder" is trying to name from the engineering side. Someone who decides what's worth making, makes it, and gets it in front of a user, without needing a committee to authorise each step. Two titles, the same converging shape. AI is chewing through one version of each role while leaving the other exposed. --- ## Which version of each role survives Both titles have always covered several different jobs under the same word. On the PM side, [Marty Cagan](https://www.svpg.com/product-management-start-here/) has been drawing this line for over a decade. There was the feature-team PM: writing tickets, refining backlogs, chasing alignment, compiling roadmaps, sitting in the middle between engineering and the business and translating between them. And there was the empowered PM: the person who sees what customers don't realise they need, decides what's worth building, defines what success looks like before the work starts, and knows when to stop. Same title. Very different jobs. On the engineering side, the same fork. [John Cutler](https://cutle.fish/blog/12-signs-youre-working-in-a-feature-factory/) named this one nearly a decade ago: the feature-factory engineer, picking up the story, shipping it, closing the loop, moving to the next one. In it for the points. And there was the product engineer, the "engineer with commercial instincts" that every hiring manager claims they want and few companies reward. The one who read the support queue before writing the feature, who pushed back on the PRD because the design didn't match how customers actually behaved, who spotted the obvious thing nobody had thought to add. The feature-team PM and the feature-factory engineer were often the default. The empowered PM and the product engineer existed, but they tended to exist in spite of the system rather than because of it. Coordination work got measured and rewarded. Tickets closed and sprints hit were visible; the decision not to build the feature that nobody actually needed was invisible. The visible version of each role paid off in visible output. The other version paid off only much later, if at all, and often by absence: the six months you didn't waste on the thing that wouldn't have worked. Nobody celebrates what didn't happen. When Rabois says PMs don't make sense, he's picturing the feature-team PM. When Cherny says the software engineer title is going away, he's picturing the feature-factory engineer. Both predictions are about the same thing: the version of each role that was mostly about running the process is being automated. The version that was always doing the real work is what's left. --- ## What agents are claiming first Cherny says Claude is already coming up with ideas on his team. It reads feedback, scans bug reports, watches telemetry for fixes worth shipping. The things that used to be a PM's Monday morning. Not in some hypothetical future, now. On the Claude Code team, the PM codes, the designer codes, the engineering manager codes, the finance guy codes, the data scientist codes. Cross-disciplinary is already the baseline, not the aspiration. Lenny asked Cherny whether the three traditional disciplines (engineering, design, product) would still persist as separate things. The answer was that in the short term yes, but there's maybe a 50% overlap in those roles already, with specialties rather than clean boundaries. The walls aren't gone. They're getting thin enough to see through. At OpenAI the prototyping stage has shifted in the same direction. [Kevin Weil's point](https://www.youtube.com/watch?v=scsW6_2SPC4), made a year ago and more true now than then, was that there's no reason to be showing stuff in Figma any more: you can vibe code a working demo in half an hour and put it in front of someone. The gap between "here's what I'm imagining" and "here's the thing, try it" has collapsed. [Anthropic launched routines](https://claude.com/blog/introducing-routines-in-claude-code) last week, joining similar features already shipped by OpenAI Codex and Cursor. Agents that run on schedules against codebases and connected systems, doing triage, reviewing, connecting dots, surfacing patterns. The coordination work that used to eat PM days is becoming infrastructure. So is a lot of the ticket-taking work that used to eat engineering days: the well-specified bug fix, the clean refactor, the boilerplate that maps one-to-one from a written requirement to working code. The spec-to-implementation step is getting very short, very fast. The temptation to skip the thinking got stronger as a result. When building was slow, the thinking stage was forced on you by the cost of getting it wrong. The friction did the work. The friction isn't there any more. --- ## New names for old skills AI labs have started naming the components of good judgement with more precision than either product management or software engineering traditionally has. Evals inherit from both sides of the divide: the PM's discipline of articulating what good looks like before the build, and the engineer's discipline of writing tests that keep checking the thing still works. Context engineering is curating the right information for a decision. Harness engineering is shaping the whole operational environment around an agent so that good decisions become the path of least resistance. [Nick Turley at OpenAI](https://www.youtube.com/watch?v=ixY2PvQJ0To) put his finger on what these names actually are. He'd been writing evals before he knew what an eval was, because he'd been outlining the ideal behaviour for a use case before the work started. His way of putting it: > "it's not that different from the wisdom of, you ought to articulate success before you do anything else" These are new names for old skills, borrowed from both sides of the divide. But AI teams are more rigorous about them than either discipline was alone. Success criteria get version-controlled instead of written once and forgotten. Feedback loops get built as infrastructure instead of left to social processes that break under deadline pressure. When an eval fails, the system tells you. When a PRD's success criteria drift, you find out in a retro six months later, if at all. If you've read my previous posts: evals are the same shape as [kill criteria](https://henty.com/blog/the-discipline-most-founders-skip), just applied to model behaviour instead of experiments. The argument for pre-committed success criteria is the same argument. And context engineering is the name AI engineering has given to [the problem I spent the last post describing](https://henty.com/blog/the-next-challenge-for-ai): context needs to flow to where decisions happen, not sit in a silo waiting to be found. An agent harness is what good management has always been trying to build: the stack of tools, context, and guardrails that make good decisions the path of least resistance and bad decisions hard to ship. The coordination overhead used to be that harness, crudely, through meetings and stand-ups and ticket conventions. What replaces it is tighter, more automated, and built around judgement rather than around keeping everyone busy. One thing to watch. In the same interview, Rabois observes that the number one consumer of tokens in the best organisations is now the CMO, and reads that as a signal of the CMO's performance. [At Nvidia's GTC a few weeks earlier](https://www.tomshardware.com/tech-industry/artificial-intelligence/jensen-huang-says-nvidia-engineers-should-use-ai-tokens-worth-half-their-annual-salary-every-year-to-be-fully-productive-compares-not-using-ai-to-using-paper-and-pencil-for-designing-chips), Jensen Huang had said he'd be "deeply alarmed" if a $500,000 engineer wasn't consuming at least $250,000 worth of tokens a year. Maybe. But "tokens consumed" is just the latest in a long line of output proxies: story points burned down, tickets closed, slides shipped. Each one starts as a rough signal of activity and gets quietly promoted into a metric that can be gamed. The old ritual dies and a new ritual grows in its place. Naming the skills properly only works if the measurement doesn't drift back into counting activity. --- ## Smaller teams, broader roles Both titles are bending. [LinkedIn has already replaced "Associate Product Manager" with "Associate Product Builder"](https://www.lennysnewsletter.com/p/why-linkedin-is-replacing-pms) in its graduate programme. There's no longer a resume: applicants send a 60-second demo of something they actually built. [Brian Chesky merged Airbnb's product management and product marketing functions](https://blog.logrocket.com/product-management/airbnb-eliminated-traditional-pm-role-now-what/) two years ago, not to kill product management but because the two roles had become one job in practice. Cherny thinks the software engineer title is next. Titles change. They've changed before. They'll change again. I was a webmaster in a previous life. The hiring side is where the shift is already visible. The traditional path into both roles ran through the coordination layer. Juniors learned the harder work by doing the ticket work long enough to see the patterns. That ramp is getting shorter on both sides, which is exactly what the LinkedIn move is responding to: select for judgement at entry, rather than hope it emerges from years of coordination work that no longer exists. The reframe also sharpens what you're hiring for. Companies that hired feature-team PMs were testing whether candidates could run the process: write the tickets, chase the alignment, keep the trains running. Companies that hired empowered PMs were already testing for judgement: show me a product decision you owned, walk me through a trade-off, tell me about the thing you killed. Same story on the engineering side. LeetCode, a timed puzzle with a known answer, is exactly the kind of problem an agent now solves in seconds. The test that matters is whether candidates can see which tickets deserved shipping, or should never have been written. If you're still interviewing for the feature-factory version, you're hiring for a job that's being automated. The uncomfortable implication of the convergence is that the lines can't stay clear at scale. If everyone is a builder with overlapping capabilities, a team of twelve becomes a free-for-all. The threshold where a team used to need splitting was somewhere around ten or twelve: past that, communication overhead compounded and you had to break the group up. With roles blurring, [the threshold moves down](https://leaddev.com/management/ai-has-us-asking-does-team-size-still-matter). A team of five or six, each with overlapping capability, starts hitting the same friction that used to kick in at twelve. Teams still need to split. They just need to split earlier. Total company size is getting reshaped too. [Block cut its workforce by 40% in February](https://fortune.com/2026/02/27/block-jack-dorsey-ceo-xyz-stock-square-4000-ai-layoffs/), citing AI. [Snap cut 16% last week](https://techcrunch.com/2026/04/15/snap-is-cutting-1000-jobs-16-of-its-workforce/), with Evan Spiegel citing smaller focused teams and AI doing more of the repetitive work. Every announcement of this kind comes with an AI justification attached. Whatever the real drivers, the interesting question is which layer will prove compressible over time. Coordination work on the PM side, ticket-taking on the engineering side, is where an agent can do most of the job. The empowered work isn't. The hiring-side signals (LinkedIn selecting for judgement at entry, Airbnb merging PM and marketing) point the same way. If the layoffs follow that same logic, the coordination layer is the one being thinned. Whether they do or not, the story is half-told. Restructuring to do today's work with fewer people is the easy first move. It's also the one that stops working as soon as a competitor runs the second move: use the same shift to do more work than was possible before. The startups showing up from the other direction are already running the second move. [Midjourney reached $200M ARR with eleven employees](https://www.uplevelai.co.uk/10-ai-unicorns-with-tiny-teams/). [Cursor hit $100M ARR in twelve months from launch](https://research.contrary.com/company/cursor) with about twenty people. They got there by never building the coordination layer and pointing everyone at the empowered work from day one. For a bigger company watching this, doing the same with less is a trap. The point isn't "less." The point is what the empowered people are freed up to do. Hiring doesn't stop. It shifts: fewer people running the process, more people who can see a problem, decide it's worth solving, and ship it. --- ## What's left is the exciting part The threat isn't that product management dies, or that software engineering dies. The threat is that the version of the role you were doing, the one with the visible process and the predictable output, is being automated. The version that survives is the one the job description kept pointing away from. That version is the hard part. It's also the best part. Knowing what to build. Knowing why. Knowing when to stop. Shaping the thing with care rather than just hitting the spec. Being the person who can say "this is the way to frame it" in a room where everyone else is still trying to figure out what to measure. The last time a floor dropped out from under what was possible was the dotcom boom, and anyone who could build was suddenly in the room. This feels like that again, with a wider door. The person who decides what's worth building, shapes it, and gets it in front of someone who cares is now the person. Broader roles, smaller teams, and a lot of work suddenly worth doing. Call it builder, call it product manager, call it whatever ends up sticking. It always was the work. The question is whether you've been doing it. --- ## The next challenge for AI: knowing what to build URL: https://henty.com/blog/the-next-challenge-for-ai Date: 2026-04-14 AI has got very good at building things. What it hasn't figured out yet is what to build, and why. That's not an intelligence problem. It's a context problem, and it's one that human teams have been wrestling with for years, with mixed results. With agents now writing code and running experiments autonomously, it's about to matter a lot more. We had three systems. Product management, engineering, support. Everyone had access to all three. In theory, an engineer picking up a ticket could open the product tool, read the customer quotes, and build something that addressed the actual pain. The support team could surface patterns that would reshape priorities: the exact language customers used, the workarounds they'd built, the features they kept asking about. In practice, people lived where their work happened. The engineers lived in the engineering tool. The product team lived in the product tool. The support team lived in the support tool. It wasn't that nobody ever crossed over. Some did. But it depended on individual curiosity, not on how the work flowed. The spec arrived, the work got done, and the context that would have changed how it got done stayed where it was. The instinct is always to fix this with process. Weekly syncs. Shared dashboards. "Everyone should check the other systems." A new checklist on top of the last checklist. It doesn't stick. People gravitate towards the system where their work happens. Everything else is overhead they'll do for a week, maybe two, and then stop. McKinsey found that employees spend [1.8 hours per day](https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-social-economy) just searching for and gathering information. The information is there. Finding it has become the job. Nobody decided to silo the teams. It's just where things naturally end up. People use the tool that does their job, and the context stays where it was created. But when context did flow, the results were obvious. An engineer who happened to read through the support tickets before picking up a feature would build it differently. Better. More aligned with what customers actually needed rather than what the spec described. "Happened to" is the key phrase. It was serendipity, not a system. --- ## The loop There's a loop that runs in every product team, whether they've named it or not. Identify an opportunity. Build something. Measure whether it worked. Feed the results back into the next decision. At small scale, this loop runs naturally. A small remote team on Slack, maybe one or two channels for everything. Everyone sees the customer quote get pasted in. Everyone catches the thread where someone shares a support pattern. Everyone knows why yesterday's decision got made, because they were in the conversation when it happened. The context is ambient. Nobody has to go looking because it's already there. Then the team grows. One channel becomes ten. Team channels appear where people outside the team feel less welcome. Private channels emerge where context is completely hidden, sometimes even from leadership. Each split makes sense in isolation. Nobody wants to be spammed with irrelevant messages. But each one is another fracture in the shared picture. At scale, each step of the loop gets handled by different people with different tools. The people identifying opportunities aren't the people building. The people building aren't the people measuring. The people measuring aren't feeding results back to the people making the next decision. The loop doesn't stop working because someone failed. It degrades because each step is missing context from the step before it. The engineer builds the feature without knowing why it was prioritised. The PM measures outcomes without knowing what trade-offs the engineer made during implementation. The results feed back into a planning process that has already moved on to the next thing. Not everything that goes through this loop needs the same level of rigour. A button change, a copy tweak: the loop still runs, it just runs quickly. Ship it, measure it, move on. Larger things need more care. Pre-committed success criteria. Proper measurement. Someone making sure the results actually inform the next decision rather than disappearing into a dashboard nobody checks. The loop is the same at every scale. The discipline scales with the stakes. [GitLab's 2025 survey](https://about.gitlab.com/press/releases/2025-11-10-gitlab-survey-reveals-the-ai-paradox/) of over 3,000 DevSecOps professionals found that 60% of organisations use five or more tools for software development. They called it the "AI Paradox": AI accelerates coding, but fragmented toolchains create new bottlenecks. Their number: seven hours per week per person lost to inefficient processes, collaboration barriers, and limited knowledge sharing across teams. That's almost a full working day, every week, spent not on building but on finding things out and getting people on the same page. --- ## What the fastest-growing company in the world figured out Amol Avasare, head of growth at Anthropic, was on [Lenny's Podcast recently](https://www.youtube.com/watch?v=k-H4nsOTuxU) talking about something that caught my attention. Anthropic's growth team has built a system where an agent identifies opportunities for growth experiments, builds the experiment, checks it against brand guidelines and quality standards, and then analyses the results. Small experiments run fast with minimal human review. Larger ones get more oversight. The thing that makes it work isn't the agent. It's that the agent has context. Brand guidelines, previous experiment results, what's been tried before, what worked and what didn't: all of it is available to the agent as part of the process. Not sitting in a separate system waiting for someone to go find it. Built into the loop. Amol was talking about growth experiments specifically, but the loop is the same for any product decision. Identify the opportunity. Build something. Measure whether it worked. Feed the results back. The principle is the same whether you're running a growth experiment at Anthropic or deciding what feature to build next at a 20-person startup. He made another observation that stuck with me. The thing that makes this hard for larger projects isn't the loop itself. It's the cross-functional coordination. Getting six people to align. His head of design put it best: "We will have AGI and it will still be impossible to get six people in a room to align." I think shared context would take a lot of the pain out of that. Not because it eliminates disagreement. People will still disagree about priorities, about trade-offs, about what matters most. But at least they'd be disagreeing from the same starting point. I'd bet that a good chunk of the friction in those conversations comes from people arguing past each other because they're each working from a different slice of the picture. That's not a disagreement. That's a failure to share information. --- ## Agents inherit the problem and amplify it Here's where it gets interesting. And urgent. Agents are now part of the team. They're writing code, running experiments, analysing data. And they have the same problem the engineers had, except worse. Even on a fully remote team, humans accumulate context over time. The Slack thread you happened to read three weeks ago. The customer call recording someone shared. The standup where someone mentioned a pattern in the support tickets. You build up a fuzzy sense of what matters, even if you can't always point to where you heard it. Agents don't have that. They have exactly what you give them. Nothing more. An agent building a feature has no idea that the support team fielded 47 tickets about the same friction last month, unless that context is somewhere the agent can reach. It doesn't accumulate a sense of things over time. It starts cold every session. The current wave of solutions looks like this: connect your agent to your project tracker. Give Cursor your codebase. Set up an integration for your support system. Each agent gets smarter about its own slice. The engineer's agent knows the codebase inside out. The PM's agent can recite every opportunity. The support agent sees every ticket. But this is the same trap. Three teams with three agents, each with deep context about one silo and virtually zero awareness of the others. The engineer's agent has never seen a customer quote. The PM's agent doesn't know what shipped last week. The support agent has no idea what's being built. A [recent report on context management](https://datahub.com/blog/context-management-strategies/) found that 57% of organisations are duplicating AI efforts across departments because there's no shared context layer connecting them. Each team builds its own context, picks its own tools, and defines its own version of what matters. Apparently we're now in the era of AI psychosis. Spend long enough talking to an agent and you come out convinced you've cracked it. The thing keeps refining your reasoning, agreeing with your approach, telling you the logic is sound. (Is that actually psychosis, or am I stretching the term to make a point? A bit of both.) Now multiply it across a team. The PM's agent confirms the PM's priorities are spot on. The engineer's agent confirms the technical approach is solid. The support lead's agent confirms the customers are being heard. Everyone walks into the meeting more certain than ever, armed with AI-backed evidence, all looking at a different slice of the same picture. Alignment just got harder, not easier. Individual tool memory is the same problem as "everyone has access to all three tools" was. The information is technically reachable. Nothing connects it. You've added agents to the team and given them the same fragmented view that was already failing the humans. The problem isn't giving agents more memory. It's giving them our shared context. We might still be deluded, but at least we'd be on the same page. --- ## A question worth asking I keep seeing this pattern at every scale. A small team where context flows naturally. A growing team where it fragments across three tools. A larger company where the fragmentation has become so deep that entire teams duplicate work because they don't know what the team next door already tried. And now, agents entering the loop with deep expertise in their own silo and virtually no awareness of anything outside it. Each one getting really good at its own job while the bigger picture drifts. The information exists. It always exists. Customer quotes, support patterns, experiment results, the reasoning behind last month's decision. It's all somewhere. The problem has never been generating context. We generate mountains of it. The problem is that it's not where it needs to be when someone, or something, is making a decision. So here's the question I think every team should be asking right now: when someone picks up a task (a teammate, an agent, you) what do they actually know? Do they know why this work was prioritised? Do they know what customers have said about this problem? Do they know what was tried before and why it didn't work? Or do they just know what's in the ticket? And the harder follow-up: if they don't know those things, how is the agent any different from the engineer who never checked the product tool? I don't think anyone has fully solved this yet. Anthropic is further along than most because they're building the tools and using them at the same time. But I think the shape of the answer is becoming clearer: context can't be something you go and find. It has to be something that flows to where decisions happen, whether those decisions are being made by a person or an agent. AI can build anything now. The next challenge is making sure it knows what's worth building. --- ## The discipline most founders skip URL: https://henty.com/blog/the-discipline-most-founders-skip Date: 2026-03-09 We had everything in place. The channel partners were lined up and briefed. The systems were ready. We'd spent months building this Wizard of Oz service where we'd do everything behind the scenes, white glove, one customer at a time, while the app pretended to handle it all. We'd rehearsed the handoffs. We'd prepared for volume. The day we went live, I remember checking my phone. Then checking it again. Refreshing the dashboard. Waiting for the first call. It didn't come. Neither did the next day's. By the end of the first week, nothing. Crickets. We were expecting to be inundated. Our channel partners had been adamant: the demand is there, the customers are ready, you need to move now. We felt real pressure to deliver. The reality was nothing like that. Not a trickle that needed time to build. Not a slow start with promising signals underneath. Just silence. But instead of asking what that silence meant, we kept banging away at the channel. Maybe we need different messaging. Maybe we need more partners. Maybe if we just tweak this one thing. There was always one more thing to tweak. Six months later we were still tweaking. Still pushing. Still explaining to ourselves why the customers were just around the corner. The whole thing should have died in week three. The writing was on the wall. We just didn't want to read it. --- ## Why this keeps happening This isn't an unusual story. If you've built anything, you'll recognise it. The problem is that building is addictive. It is so easy to get swept away by the conviction that you've got something everybody's going to want. Easy to get pulled and seduced by the need to add just one more thing. Before you know it, the lost days turn into weeks, then into months. And when the results don't come, that same conviction turns into explanation. We've invested too much to stop now. If we just tweak the approach, the customers will come. There's always one more thing to try. It only takes one person who hasn't committed to what "enough" looks like, and every week the goalposts move a little further. Nobody has to admit they've moved. The pattern is always the same. Intuition over evidence. Conviction over curiosity. Not that intuition and conviction are wrong. They're essential. But they need to be balanced with evidence and curiosity, and most of the time they aren't. [Teresa Torres](https://www.producttalk.org/), who wrote *Continuous Discovery Habits*, puts her finger on one version of this: teams interview customers to confirm what they already believe, rather than to discover what they don't know. The research happens. The conversations happen. But the question going in is "am I right?" not "what don't I know?" And so the evidence gets shaped to fit the conviction, rather than the other way around. --- ## What Disciplined Entrepreneurship gets right Bill Aulet's *Disciplined Entrepreneurship* is the best antidote I've found to this pattern. The book lays out 24 steps for building a startup, but the principle underneath all of them is simple: the undisciplined founder wastes time being optimistic about the wrong thing. The disciplined founder earns the right to be optimistic because they've verified something real. That's not pessimism. It's just not wasting time. The concept that sticks with me most is the beachhead market. Your market is not "everyone." You cannot validate against "everyone." A beachhead market is a deliberately small, winnable segment. The one place you plant your flag first and take completely before you expand. Choosing a beachhead forces you to be specific about who the customer actually is, what they actually need, and what a real win looks like. Without that specificity, no amount of customer research will give you a clear signal. You'll just be collecting noise and calling it data. Aulet is also explicit about something most founders would rather skip: there's no substitute for talking to actual customers. Reports, desk research, analyst opinions: none of it counts. Direct contact. Interviews. Observation. Prototype testing. The customer is the most important element of the entire framework, and all 24 steps begin and end with understanding a specific customer in a specific context. --- ## Kill criteria: the discipline that changes everything Here's the thing that would have saved us six months on that experiment: kill criteria written before we started. Kill criteria are the conditions under which you stop. Not "we'll see how it goes." Not "if we don't see traction." Specific, measurable, pre-committed conditions. If X doesn't happen by date Y, we stop. No renegotiation. The reason they have to be written before the experiment starts is the same reason Ulysses had his men tie him to the mast before the ship reached the Sirens. He knew that once he could hear the song, he would want to change course. He would beg, plead, rationalise. So he pre-committed. He bound himself while his judgment was still clear, precisely because he knew his future self would try to wriggle free. That's the Ulysses contract. And it's exactly what kill criteria are: a pre-commitment made when you have clear judgment, designed to bind your future self when your judgment is impaired by sunk cost and emotional investment. Part of what makes this hard is that you can never truly prove nobody wants what you're building. There's always another segment to try, another message to test, another door to knock on. That ambiguity is the trap. Because if you can never prove a negative, you can always keep going. Kill criteria don't solve that philosophical problem. What they do is decide in advance how much evidence of absence is enough. You're not trying to prove nobody wants it. You're agreeing on what "enough signal to stop" looks like, before you're too invested to see it clearly. Without kill criteria, what happens is predictable. The experiment underperforms. The team discusses it. Someone says "we should give it more time." Someone else says "maybe we should tweak the approach." Nobody says "we should stop," because stopping feels like failure and nobody defined what failure actually looks like. So the goalposts move, the timeline extends, and six months later you're still banging on doors that nobody is answering. With kill criteria, the conversation changes completely. "We said we'd need 50 sign-ups by March. We have 3. We stop." There's nothing to renegotiate. The decision was made when the thinking was clear. This matters just as much if you're building alone as it does if you have co-founders. Solo, kill criteria are a commitment to yourself, a way of holding your future self accountable when conviction starts to cloud the evidence. With co-founders, they do something equally valuable: they create alignment before the pressure arrives. Everyone agreed on the conditions when nobody was emotionally invested. When the moment comes, there's no ambiguity about what you all said you'd do. If you're on a product team, you'll notice that kill criteria look a lot like key results. They should. The best OKRs measure outcomes, not outputs, and kill criteria are the same thing applied to experiments: not "did we ship it?" but "did anyone care?" | Vague criteria | Kill criteria | |---|---| | "If the numbers look good we'll continue" | "50 sign-ups by March 15 or we stop" | | "We'll see if there's traction" | "3 paying customers in 60 days or we pivot" | | "If customers seem interested" | "Fewer than 10% of pilot users return in week 2: kill the feature" | *The left column gives you permission to keep going forever. The right column forces a decision.* --- ## The evidence hierarchy Not all signal is equal. This is something I wish I'd understood more clearly earlier. The weakest signal is explicit demand. Someone telling you "I would use this." People say that all the time. It costs them nothing to say it. Surveys, focus groups, interviews where someone says "yes, I'd definitely buy that": this is the least reliable data you can collect. It feels good. It feels like validation. It isn't. Go back to that opening story. Our channel partners were adamant. The demand is there, the customers are ready. But that was their read of their customers, secondhand and filtered through their own interests. They weren't lying. They believed it. But belief isn't the same as evidence, and their customers' stated intentions weren't the same as actual demand. When we opened the doors, the customers didn't show up. The channel partners' conviction was no substitute for going and finding out directly. Stronger signal comes from behaviour. Clicks. Sign-ups. Payments. Someone actually doing the thing, not just saying they would. The gap between what people say they want and what they actually do is enormous, and most founders are gathering the wrong type of signal. The strongest signal, and the hardest to see, is latent demand. Adjacent frustration. Compensating behaviours. Workarounds that people have built because the thing they actually need doesn't exist yet. The iPhone is the most cited example of this, and with good reason. Before 2007, nobody was asking for a device that combined a phone, an iPod, and an internet browser. The explicit demand wasn't there. But the latent signal was everywhere: people carrying three devices. People hacking their phones to do things phones weren't designed to do. The frustration was real. It was just expressed through behaviour, not words. Apple read that signal. Behavioural rather than stated preference. They looked at what people were doing, not what people said they wanted. Study after study finds that the majority of shipped features are rarely or never used. The [exact figure is debatable](https://www.mountaingoatsoftware.com/blog/are-64-of-features-really-rarely-or-never-used), the [pattern is not](https://www.pendo.io/resources/the-2019-feature-adoption-report/). Most of what gets built didn't need to be built. The signal was there; it just wasn't the signal anyone was looking for. --- ## The loop, not the gate The most common mental model mistake is treating validation as a phase. Something you do before you build, a gate you pass through once and then you're free. I used to think that way too. Traditionally, the build stage would be delayed as long as possible until we could really validate. And that made sense when building was slow and expensive. Here's the counterintuitive part: now, with vibe coding and rapid prototyping, the gap between discovery and build is collapsing. You can get something into the hands of users in days, not months. And it's tempting to think: what have I got to lose? Just build it, see if people want it, move on. The problem is that "just build it" is still a bet. You build the thing in a few days, discover nobody wants it, build another thing, discover nobody wants that either. Before you know it, six months have gone and maybe one of them worked. You're in exactly the same place as if you'd spent six months building the wrong thing the slow way. The only difference is that you mistook activity for progress the whole time. The bottleneck is no longer the building. The bottleneck is the validation. Speed doesn't change that. It just makes it easier to avoid facing it. Both things are true at the same time. Build fast. Validate fast. Get it into the hands of users as quickly as possible. Talk to users as much as possible. But the loop is still there; it's just got tighter. The discipline isn't in delaying the build. The discipline is in staying honest about what the evidence is telling you while you build. Building without that honesty is just moving faster in the wrong direction. Torres calls this continuous discovery: the idea that discovery is not a phase but a weekly habit, something teams do alongside building, not before it. Her [opportunity solution tree](https://www.producttalk.org/opportunity-solution-trees/) forces you to name the assumption you're testing and, if the assumption is wrong, prune the branch. You don't renegotiate; you move on. The validation loop never ends. Every feature request starts with a problem worth solving. Every pivot goes back to the beginning. Every new market segment means asking the same questions again. The loop doesn't stop when you launch. It doesn't stop when you find product-market fit. It just gets tighter. --- ## What this means for how you work Before your next experiment, three things. Write kill criteria before you start. Specific conditions, measurable outcomes, a date. "If X hasn't happened by Y, we stop." Write them down. Share them with the team. And when the results come in, let them inform the next experiment, not rewrite the rules of this one. Question what type of signal you're collecting. If your validation is based on people telling you they'd use something, that's the weakest signal available. Look for behaviour instead. Look for workarounds. Look for adjacent frustration, the things people are already doing that tell you there's a real problem underneath. Treat validation as a loop, not a gate. You're never done. The question is never "have we validated?" It's "what are we validating right now?" Every feature, every experiment, every new direction goes back to the same question: is there a real problem worth solving? Building is more fun than validating. It always will be. And now that building is faster and easier than ever, the temptation to skip the discipline is stronger than ever too. But the skill of defining exactly the core value of what you're offering and testing whether anyone actually wants it: that skill hasn't become less important. If anything, it's become more important, because now you can waste time faster than ever before. The discipline most founders skip isn't complicated. It's just uncomfortable. It means writing down the conditions under which you'll stop before you've fallen in love with the idea. It means looking for evidence that you're wrong, not just evidence that you're right. It means treating silence as an answer, not as a problem to be solved with more noise. That experiment we ran? The one that should have died in week three? The evidence was there. We just weren't looking for it. Or maybe we were, but we didn't want to see what it was telling us. Next time, I'll tie myself to the mast first.