Triage separates simple changes from work needing specifications
Original excerpt
If this is easy and this is unambiguous, just implement it,
Additional context is not included in this excerpt. Read the original conversation.
THE ORIGINAL CONVERSATION
AI Engineer · · 20:36
Zach Lloyd outlines a software-factory workflow: triage, specification, review, and verification. He also discusses using observer agents to improve how other agents apply skills.

Selected viewpoints with original excerpts and context. Read the full transcript below.
Original excerpt
If this is easy and this is unambiguous, just implement it,
Additional context is not included in this excerpt. Read the original conversation.
Original excerpt
I would have an agent do code review first, and then it becomes over time, like, a risk management exercise
Additional context is not included in this excerpt. Read the original conversation.
Original excerpt
a common kinda loop that you're gonna wanna put in your factory is like a skill loop. That means you're gonna have your factory agents that are running skills, and then you'll have observer agents that are seeing how those skills are being applied, looking for issues, and trying to improve the skills.
Additional context is not included in this excerpt. Read the original conversation.
Publisher timestamps; playback alignment is pending.
This publisher provides subtitle timings without chapter markers.
Publisher transcript in English.
AI review covers the selected viewpoints. The publisher transcript is preserved as supplied; full-text and playback checks are pending.
Okay. Hello, everyone. Uh, I'm excited to be here. Uh, my name is Zach Lloyd. Um, today I'm gonna be talking about self-improving software factories, the new open source model, and basically what I think is happening to development. Little bit about me just to begin. So I am a former, uh, principal engineer from Google, basically lead engineering on the Google Docs suite. I've been an engineer now for over twenty years. Long time. I am still, uh, shipping frequently, but I haven't written a line of code in the last six months. Uh, and I'm the founder of a company called Warp. Uh, Warp, if you're not familiar, is a open source agentic development environment. Uh, you may know us as a terminal. That is how the company started. We're basically a terminal that has agents built in. Uh, we open sourced it a couple months ago, and I'm gonna talk a little bit about that experience and the motivation for it. It's a popular open source project, over sixty thousand GitHub
stars. We've had a couple hundred people contributing. We have over eight hundred thousand active developers who are using Warp. And increasingly, we are focused not just on the terminal aspect and the interactive aspect of development, but more so on how do you automate development. I'm gonna talk mostly about that. So the thesis that I have, uh, is that the discipline of software engineering is going to become something more like factory engineering, and I'll explain what I mean by this in a minute, but just keep that in mind. That's, that's what I think is gonna happen. Um, if you look at development over the past couple years, it's-- I mean, it's just crazy how it's changed. We've gone from a world of chat and AI auto-complete, so Cursor, Copilot, to the phase that we're in now, which I consider to be mostly interactive agents. So you're sorta sitting at your computer, and you
are telling Claude Code to do something. You're telling Warp to do something. And I believe what's gonna happen over the next six months, a year, hard to predict the pace, is that we're gonna move much more towards a world of automation. But before I get into that, just quick show of hands, how many folks in here are building with agents? Ev-every... A hundred percent. Makes sense. Uh, how many, how many folks are building typically with m-multiple agents at one time? So again, almost everyone. How many people are running an agent right now? I'm not offended. I would-- Okay, that's totally cool. I, I would be doing it too. Uh, how many folks are running agents in the cloud, out of curiosity?
So that looks like less than half, but still significant. Uh, and how many folks have, have set up a system internally to automate the whole software development life cycle? So everything from, like, triaging, speccing, implementing, reviewing. So I see some hands. So some people are doing this. So this is what's gonna happen. Uh, every project of significant size, I believe, is gonna have something like this. Um, and it's gonna look kinda like this big loop, and everyone is talking about loops. There's nothing that complicated about loops. This loop says a cloud software factory. This loop could literally just say, like, the software development life cycle. It's the same thing. Um, but just to go through this loop, it's like ideas are gonna come in at the top. Agents are gonna do triage. If something is complicated, they will write a spec. Uh, these little blue boxes are where humans step in. Humans will review the spec. Uh, agents will do the implementation.
A human and agent will review the code. Agents will verify. Human will review the product. You ship, and then you monitor, and round and round you go. Uh, and this is what software development, for better or worse, I think is gonna end up looking like. So I repeat the thesis, which is that if this is what software engineering is gonna look like, um, software engineers are going to be the ones who end up building and managing these factories. Now, I promised at the beginning, and I put in the title of the talk that I was gonna talk about open source, and so I wanna do that for a few minutes. I'm gonna take a quick digression. Um, I bring up open source because one of the main reasons that Warp open sourced was to build, uh, build a public factory. Um, and so this is a picture of this website we've built called build.warp.dev, which shows all of the issues that are flowing through our system and what state they're in, what agents are working on them, what contributors are working on them, and it's
kinda like a proto-factory done at scale. It's not working perfectly, but it is working. And one of the reasons we open sourced was to try to build this. Um, just in general, I think it's interesting to talk about open source in the time of agentic development. This is a really stupid graph, but it's like... It just-- It-- You get it. It's like, it's becoming much cheaper to build software. Um, a corollary of that is that it's becoming trivial to clone software. And so if you are in the software business, and I don't know how many folks in this room are in the software business per se, but it's very hard to build a software business if it's free to build software. It's hard to capture the value, especially if a competitor can clone. And so my big takeaway or big tip for everyone in here is that the first thing you should do is patent your code. Uh, I'm kidding. Don't. This is a complete joke. Don't, don't do this.
Uh, my, my first tip is obviously you need to have a great product. Uh, this has always been the case, um, but I would say a great product- Probably was never enough. But even now, more than ever, if you think that you're gonna build a great software business just by building and shipping a great product, you're probably n- not gonna succeed. Uh, you need advantages beyond the product. And so, those advantages could look like distribution, ecosystem, it could be that you have a great brand or a data moat. You might have capital, um, but if you're a startup, again, I'm coming from the startup world here, you just don't have these advantages. And so, you, you still wanna break through, uh, and one of the ways that I suggest doing this is by building in the open. And so, to be clear, it took Warp five years of building closed to sort of make the leap into building in the open, and I'll explain why. But if you build in the open, um, it helps build your ecosystem.
It, it can take you from being, like, hated on Hacker News to, like, tolerated. Uh, it can burnish your brand, it creates community. And so, there's all these advantages to it. Uh, and some of the things that I think have traditionally been a pain, um, can now be managed. And so, like, the traditional, you know, pain of open source might be something like, you get a lot of noisy issues. You get sloppy PRs. You can end up in code review hell. You can end up having to spend a lot of time verifying changes. And so, the solution, so kind of a long-winded way of getting to this, for open source or at least for Warp in the case of open source, the thing that made us finally decide to do this, was that we built a whole set of automations, really a software factory, around managing the open source project. And so, like I said, this is what I think the future is gonna look like. I'm gonna drill into it a bit just to get a little bit more technical for folks who want to try to build
something like this for their own projects. So, what are the components of an effective software factory? Uh, it's really not that complicated to start or at a high level. You need a set of automations. You need a way of providing context and skills. You need a way of bringing humans in at the correct time to sort of like when things get stuck on the factory. And then a really important thing is you need some set of self-improvement capabilities. So, think of this as loops. And if you do this right, uh, in the open source world, you can get a, you know, world where agents are helping contributors contribute, they're helping maintainers maintain. And I wanna emphasize, there's nothing special about open source here. I think every sizable project can benefit from this approach, and I, I predict that every, every company, every open source project will have at its core a software factory, kind of like the way that CI/CD became just like, "Oh, of course you have that."
Uh, maybe, I don't know when that happened, ten years ago. Let's tour the tac-- uh, the factory floor for a second here. So, you're not gonna look at this. This is too much. Uh, the, the point of this slide is not to have you read the workflow. It's that the factory floor is basically a graph, um, of steps where you are defining, like, okay, how does software get built for my product? And it looks pretty similar for every product. Things come in, they flow through, um, they get stuck at certain points. Um, and, you know, broadly speaking, just to back out a second, so there's the inputs. The inputs are really ideas. Um, the inputs could be coming from your team, they could be coming from your users. Um, the inputs themselves tend to come in through certain channels that you should think about as, like, your task tracker is an obvious one, or Slack, your teams, like, your communication channels. It could come directly from, like, your terminal or IDE.
They could come from your monitoring systems. But there's some set of inputs that bring work into the factory. There's triage. This is a really important step. So, again, I boil it down to something very simple, but, like, you want an agent that is looking at issues as they come in, and just saying, "You know what? If this is easy and this is unambiguous, just implement it," and this is how you can actually get going with a factory. If an issue is hard, uh, I recommend having, uh, an agent that produces specs. Folks in here using spec-driven development? Show of hands. People follow this? Okay. You can do this many different ways, but I think it's very effective. The way that, uh, we do it at Warp that I recommend is having an agent write what we call a product spec and a tech spec. Product spec describes the product invariance that you're building towards. Tech spec describes the architecture and the shape of the code. Then you have an implementation agent. This is basically
a coding agent that runs somewhere in the cloud. It makes a diff. You can use all sorts of coding agents for this. You have review. This is, in many ways, the most painful part. Like, I expect that people are a little bit tired of reviewing agentic slop. Uh, I would have an agent do code review first, and then it becomes over time, like, a risk management exercise of, like, when do you bring in humans to do code review? But you wanna have a step in here where humans can do it. This is a very important step, uh, for certain types of apps, the verification step. So, this would be things like computer use, uh, if you're building a, a sort of UI, having the computer actually use the code that the agent produced and producing videos and screenshots. CI/CD, still use it obviously. Uh, and then monitoring. So agents don't stop in your factory when code is shipped. They should observe what's been shipped. Is it crashing? Is it being used? And
round and round you go, 'cause you take the output of this monitoring step and you feed it back into the top of the factory. Now, you could try and build this and, um, I went to a talk earlier that my friend Adam gave, where Uber has built an internal version of this, um, and it's pretty cool. Um- I would say for most, you know, most organizations, it really depends where you are, you'll be able to build a simple version of this easily. But to build a thing that actually scales is probably-- Like, you should probably be focusing on your own product, not building this infrastructure, because there's a lot of stuff that you end up wanting. You're not-- Again, you're not supposed to read this. Uh, it's just a lot of stuff. If you do build it, um, or if you buy it, you'll end up with something that looks kinda like this, which is, um, you're gonna have a bunch of ways of getting work into your factory. That's what's at the top here. You're gonna have a sorta control plane for figuring out how work gets distributed across your factory floor. You're gonna have the actual
place where the work happens, and so that's gonna be cloud sandboxes. It's gonna be figuring out what agent to run, so what's the harness, what's the model. And then finally, I think this is a really important thing, you're gonna wanna set up some kind of data plane that sits below your factory. And so that's something that lets agents remember what they've done, learn, um, improve over time. The factory is not just like a product, it's also a mindset, and this brings me back to the thesis I had at the beginning. You need to measure and improve. So factory, this is where the-- I don't know, you can stretch this metaphor as far as you want, but like you should be thinking of efficiency. And so that means like how much software did you ship, how much did it cost in terms of human time and token time, and you're gonna wanna measure this and
try to improve it over time. A key part of this is creating loops. So, uh, loops are, again, they, they sound complicated. They're not that complicated. Loops are basically ways of, uh, having agents improve, um, by like observing what they're doing, where they're failing. Um, and so a common kinda loop that you're gonna wanna put in your factory is like a skill loop. That means you're gonna have your factory agents that are running skills, and then you'll have observer agents that are seeing how those skills are being applied, looking for issues, and trying to improve the skills. So for instance, if you had a code review agent, uh, and it was leaving comments, and a senior engineer on your team was going and correcting those comments, you'd want an observer agent that would look at that and basically, uh, improve the code review agent for the next run.
This is one thought just to leave folks with, like what is-- where does this leave engineers? Um, I think you're gonna have to get into this mindset, and I'm trying very hard to get our team into this mindset, it's not always easy, that you're not just building the product, but you're building the thing that builds the product. And that's like, it's just different. It's more like process engineering or manufacturing or something like that. Um, and you could think like, "Okay, maybe that's a bummer." Like, is that a bummer? Is that, uh, you know... And it depends. Like, it depends what joy you get out of software engineering. If your joy is in writing the code, I think you're-- everyone in, in here who is a software engineer is going to be writing less code. But if your joy is in shipping product, like it's never been a better time, and this is actually where I find my joy. It's like I like building and shipping the thing. So everyone in here is gonna code less, but they're gonna ship more, and that's gonna be a
trade-off. But if you approach it like you're a, a factory engineer, I think you can see like that there's still a really cool set of engineering challenges. Um, you could almost think of it as like meta engineering. Like how do you engineer your system of agents to be the best possible at engineering? But I think it's a very compelling and interesting set of challenges to solve. So that's it. Um, uh, for folks who are interested, uh, this-- if you follow the link on this QR code, I've set up, um, a open source GitHub repo where anyone who wants to try building their own factory agents can do it. This uses Warp's, uh, agent platform as part of it, but you honestly, you don't have to use it. I'm not trying to like push into our product. But this should give you a good sense of like, okay, if you wanna set up, uh, an agent that does, uh, triage or an agent that does, uh, spec writing, how do you actually do that? How do you get from
like the theory of, uh, working with a factory to actually putting it into practice? Um, I don't know if we have the capability to do questions in here. I saved a few minutes for questions if anyone has questions. Otherwise, I will, I will wrap up. Yes. Yeah, I have a question. So there's a bit of a tension in what you said where it's like you don't want to build this because it's a lot of work. Yes. But you are also a factory builder. Like, where do you stand on that? It's a great question. So I said something that's almost contradictory. I think, um, you should-- The way you should think of it is like y- everyone's gonna deploy some sort of factory, but then the tuning of the factory, the like are these the right skills for my domain? Is this factory building my product in the right way? I still think there's a bunch of interesting engineering challenges. And for some places you can build this. But a-again, I, I think, I think that the-- you should probably be focusing on building the core product for your company for the most part. But there's a bunch of like tuning and like, uh, figuring out how to make the
factory work for your product that matters. That's a great question. Yes. Hey, Zach. Thank you so much. Um, I had a question. So if you were a college student right now- Yes ... uh, graduating and entering the workforce, where would you be spending your time? Yeah, so the, the question in case people couldn't hear was like if I was a college student graduating and entering the workforce right now. So I think that the most important skills in this new world are adaptability. I think that that's critical thinking. It's like the s- the speed at which you can learn. I do think, I don't know if you're a computer science student, but I still think there's a ton of value in understanding like the underlying systems and architecture, and being able to reason and understand. The code,
understand the specs that agents are written, sorry, are writing. So I, I would focus on, on those skills. Um, and like, I don't know, we're hiring more people than we've ever hired. There's a lot of kind of like misdirection around like, you know, people not being hired because of AI. That's not the experience we've had so far. And what I'm looking for are, like, really adaptable, product-focused thinkers who can, like, basically be great problem solvent- problem solvers, uh, even as the underlying technology changes. Thank you. Yeah. I think I have time for one more question. Yes. What about the product factory, discovery, ideation, design, vision, taste, how about those kinds of things? What about the... Like, so the question was, what about the, the product taste in like the... How do you actually build something useful, I think is probably the right, like, maybe the framing and like, uh, where do the ideas come from? And so I think that, um, the problem with the factory metaphor, even though I'm like leaning into it because I think
that's like, that there's something to it, is that it can kinda sound like, uh, mechanizing or dehumanizing. Um, and I still think underlying all of this, the only thing that matters is like, are you building something useful? And if you have like a, a factory that is like churning out shit that no one cares about, it's like, what's the point? And I think that h-human taste, human input, human product sense, um, humans like guiding at those touch points where you can't automate stuff is absolutely, like, essential and like, that's like what I do. Um, like I'm trying to figure out what do, what do customers want, what do people want, what's gonna be valuable to them? So I think that's an absolutely key point to it. I think I'm at time, so I, I have to go. I hope folks, uh, enjoyed this chat. Uh, I'm really grateful for being invited to speak. So thank you all very much.
Zach Lloyd introduces himself as Warp's founder and a former Google principal engineer who led engineering on Google Docs. He argues that software engineering is moving from interactive coding agents toward automated software factories, with engineers increasingly designing and managing the systems that build products. His proposed loop combines agent triage, specifications, implementation, review, verification and monitoring with human checkpoints. He says Warp's move to open source supported a public factory and ecosystem, while automation helped manage noisy issues, poor pull requests and verification work. For harder tasks, he recommends separate product and technical specifications; for review and verification, he describes agent-first code review, human intervention based on risk, computer-use checks and continued CI/CD. The underlying architecture includes work intake, a control plane, cloud sandboxes, agent harnesses and a data plane for memory and improvement. Observer agents can improve skills using failures and human corrections, while efficiency should be measured through software shipped, human time and token costs. He introduces an unnamed starter repository and advises organizations to focus on their core products while tuning factory skills to their domains. In Q&A, he emphasizes adaptability, critical thinking, systems understanding and human product taste as essential to building useful software.
Editorial summaries and translations are generated from this publisher material.
This page uses an offline-processed transcript and translated reading material. Check the original recording for exact wording and timing.
Open transcript or source material (opens in a new tab)Report an issue