# Entrevista a Boris Cherny (creador de Claude Code) + Cat (PM)

**Fuente:** video del post de X de [@0xeduu](https://x.com/0xeduu/status/2073821838833070191) (05-jul-2026).
**Contenido real:** entrevista/podcast (inglés) a **Boris Cherny** y **Cat** del equipo de Claude Code en Anthropic, con intro/outro en español montado por el creador del tweet.
**Duración:** ~75 min · **Transcripción:** faster-whisper (base, es/en) · 15.309 palabras.

> ⚠️ El gancho del tweet ("casi todos los ingenieros de Anthropic corren +100 agentes a la vez") **no es una cita literal** del video. Lo que sí dicen: todos lo usan a diario (DAU "vertical por días"), subagentes que investigan en paralelo (3-5 intentos) y eligen/summarizan la mejor opción, y productividad anecdótica de 2X-10X.

---

## Resumen ejecutivo

### 1. Qué es Claude Code y de dónde salió
- Corre en la **terminal** (no web/desktop) → acceso a bash, a todos los archivos del directorio, operando de forma **agéntica**.
- Nació como **experimento personal de Boris** sobre la API pública (al inicio reaccionaba a su música / describía su reproductor). Al darle terminal + capacidad de codear, "de repente se volvió útil".
- Core team → todos los ingenieros/researchers de Anthropic → **gráfico DAU "vertical por días"** = señal de PMF → apertura al público. Equipo original: Boris, Sid, Ben. Cat entró como PM.

### 2. Filosofía de producto
- Principio central: **"do the simple thing first"** — staff mínimo, mantenerlo crudo, las restricciones ayudan.
- **Compact** = pedirle a Claude que resuma los mensajes previos (lo simple ganó a lo complejo).
- **CLAUDE.md** = solo un archivo de contexto (descartaron arquitecturas de memoria elaboradas).
- **3 capas para construir:** (a) en el modelo, (b) scaffolding sobre Claude Code, (c) Claude Code como pieza de un workflow mayor (ej. tmux, sesiones paralelas). No meter en el harness lo que vive fuera.
- Roadmap **bottom-up**: cada quien construye la feature que desearía tener; casi nada top-down.
- Diseñan pensando en **"qué sabrá hacer el modelo en 3 meses"** (más autonomía, explorar/encontrar info, exhaustividad, componer herramientas).

### 3. Anécdotas técnicas
- **La noche antes del lanzamiento**, ninguna librería de markdown renderizaba bien en terminal → Boris le pidió a Claude un **parser de markdown desde cero**, listo en 1-2 prompts. Ese es el que hoy hace ver bonito el markdown.
- **Reescribieron el harness ~5 veces**, re-simplificándolo cada 3-4 semanas. Mantener el harness minimalista para que **el contexto llegue puro al modelo** y el andamiaje no interfiera con la intención del usuario.
- Lección: baja el costo de escribir código → ahora conviene escribir tu propia librería/feature en vez de tolerar una que no te gusta.

### 4. Adopción no-técnica y productividad
- **Megan (diseñadora, no codea)** mergea PRs a la consola; data scientists lo usan para BigQuery; finanzas lo usa con CSVs (`cat archivo.csv | claude`).
- Productividad **anecdótica**: Boris ~2X; algunos ingenieros ~10X; uso mínimo (commit messages) ~10%. Aún sin número oficial.
- Debate sobre dónde entra el técnico en el flujo del no-técnico. Hoy **prompear bien importa**; "the better lesson always wins" → quizá en meses ya no haga falta.

### 5. Producto / futuro
- **Equipo permanente** creciendo. Modelo **pay-as-you-go** (sin compromiso inicial, encaja con scriptear Claude Code de forma autónoma); evalúan suscripción por predictibilidad.
- **Enterprise:** multiplicador de productividad para ICs; acompañan dudas de seguridad/monitoreo.
- **Open source:** "lo estamos investigando", aún no. El "secreto" está en el **modelo**, no en el harness (JS inspeccionable, deliberadamente minimalista).

---

## Principios transferibles (para contrastar contra otras plataformas agénticas)

1. **Do the simple thing first** — la solución mínima que funciona gana; las restricciones son features.
2. **Harness minimalista** — reescribir para simplificar, no para acumular. El andamiaje no debe ensuciar el contexto del modelo.
3. **Contexto puro al modelo** — remover cosas que confundan al modelo; dejar que el modelo piense.
4. **3 capas de decisión** (modelo / scaffolding / composición externa) — saber en qué capa vive cada feature.
5. **Diseñar para el modelo de dentro de 3 meses**, no para el de hoy — apuntar a más autonomía.
6. **Bottom-up / dogfooding** — construir la feature que uno mismo desearía; medir adopción diaria real (DAU).
7. **Subagentes en paralelo** — lanzar N intentos, elegir/summarizar el mejor.
8. **El modelo escribe sus propias herramientas** cuando la librería externa no alcanza.

---

## Transcripción completa (cruda)

> Transcrita automáticamente. El tramo final incluye ruido de doblaje ES/EN del creador del tweet (marcado por Whisper con baja fidelidad); el contenido en inglés es el fiel.

Yo no había idea, yo no sé qué, que eran los estudios que podría hacer este tipo de cosas, me dieron todo el ingenieros y los estudios que era un poco de acceso, y casi todo el mundo fue usando todo el día, y es decir, cuando se ve una tasa, es como un subclad, es un subligion que hace esto, y cuando hago algo en Harry, yo os voy a investigar 3 veces o 5 veces o cada vez en muchos veces en Prowell, y luego el quadrón va a ser el mejor opción y se summariza eso. Gracias por habernos. Gracias. You and I know each other from before. I just realized DAGster, as well. Yeah. And then index ventures, and now, and traffic. Exactly. It's so cool to see a friend that you know from before, like now working in a topic and shipping really cool stuff. And Boris, you are a celebrity because we were just having you outside just getting coffee and people would recognize you from your video. Oh, well. Right? That's new, wasn't it? Was that neat? Yeah, I had that experience once or twice. In the last few weeks. It was surprising. Yeah. Well, thank you for making the time. We're here to talk about Cloud Code. Most people probably have heard of it. We think quite a few people have tried it. But let's get a Chris up front's definition. What is Cloud Code? Yeah, so Cloud Code is in the terminal. So, you know, Cloud has a bunch of different interfaces. There is desktop, there's web, and yeah, Cloud Code, it runs in your terminal. Because it runs in the terminal, it has access to a bunch of stuff that you just don't get if you're running on the web or on desktop or whatever. So it can run bash commands. They can see all of the files and the current directory. And it does all of that agentically. And yeah, I guess maybe it comes back to the question under the question is, where did this idea come from? And yeah, part of it was we just want to learn how Cloud, we want to learn how people use agents. We are doing this with the CLI form factor because coding is kind of a natural place where people use agents today and, you know, there's kind of product market fit for this thing. But yeah, it's just sort of this crazy research project. And obviously it's kind of bare bones and simple. But yeah, it's like an agent in your terminal. That's how the best stuff starts. Yeah, how did it start? Did you have a master plan to build Cloud Code? Or there's no master plan. When I joined Anthropic, I was experimenting with different ways to use the model kind of in different places. And the way I was doing that was through the public API, the same API that everyone else has access to. And one of the really weird experiments was this quad that runs in a terminal. And I was using it for kind of weird stuff. I was using it to look at what music I was listening to and react to that. And then, you know, like screenshot my video player and explain what's happening there and things like this. And this was like kind of a pretty quick thing to build. And it was pretty fun to play around with. And then at some point I gave it access to the terminal and the ability to code. And suddenly, it just felt very useful. I was using this thing every day. It kind of expanded from there. We gave the core team access. And they all started using it every day, which was pretty surprising. And then we gave all the engineers and researchers at Anthropic Access. And pretty soon, everyone was using it every day. And I remember we had this DAU chart for internal users. And I was just watching it. And it was vertical for days. And we're like, all right, there's something here. We got to give this to external people. So everyone else can try this, too. Yeah, that's where it came from. And were you also working with Boris a review? Or did this come out and then it started growing? And then you're like, OK, we need to maybe make this a team, so to speak? Yeah, the original team was Boris, Sid, and Ben. And over time, as more people were adopting the tool, we felt like, OK, we really have to invest in supporting it because all our researchers are using it. And this is our one lever to make them really productive. And so at that point, I was using hot code to build some visualizations. I was analyzing much in data. And sometimes it's super useful to spin up a streamlet and see all the aggregate stats at once. And hot code made it really, really easy to do. So I think I sent Boris a bunch of feedback. At some point, Boris was like, do you want to just work on this? And so that's how it happened. It was actually a little more than that on my side. You were sending all this feedback. And at the same time, we were looking for a PM. And we were looking at a few people. And then I remember telling the manager, like, hey, I want cat. That's a lot. So we're dealing with all these. I'm sure people are curious. What's the process within Anthropics to like, graduate one of these projects? So you have kind of like the a lot of girls, then you get a PM. When did you decide, OK, we should, like, it's ready to be opened up? Generally, at Anthropics, we have this product principle of do the simple thing first. And I think that the way we build product is really based on that principle. So you kind of staff things as little as you can and keep things as crappy as you can, because the constraints are actually pretty helpful. And for this case, we wanted to see some signs of product market fit before we scaled it. Yeah, I imagine. So like, we're putting on the MCP episode this week. And I imagine MCP also now has a team around it in a much the same way. It is now very much officially, like, sort of like an anthropic product. So I'm kind of curious for Cat, like, how do you view PMing something like this? Like, what is, I guess you're, like, sort of grooming the roadmap you're listening to users. And the velocity is something I've never seen coming out of Anthropics. I think I've come with a pretty light touch. I think Boris and the team are like extremely strong product thinkers. And for the vast majority of the features on our roadmap, it's actually just like people building the thing that they wish that the product had. So very little actually is top-stown. I feel like I'm mainly there to clear the path if anything gets in the way and just make sure that we're all good to go from, like, illegal marketing, et cetera, perspective. Yeah. And then I think, like, in terms of very broad roadmap or, like, long-term roadmap, I think the whole team comes together and just thinks about, OK, what do we think models will be really good at in three months? And like, let's just make sure that what we're building is really compatible with, like, the future of what models are capable of. I'd be interested to double-click on this. What will models be good at in three months? Because I think that's something that people always say to think about when building your products. But nobody knows how to think about it because everyone's just like, it's generically getting better all the time. We're getting a GI soon, so don't bother, you know, like, how do you calibrate three months of progress? I think if you look back historically, we tend to shift models every couple of months or so. So three months is just like an arbitrary number that I picked. I think the direction that we want our models to go in is being able to accomplish more and more complex tasks with as much autonomy as possible. And so this includes things like making sure the models are able to explore and find the right information that they need to accomplish a task, making sure that models are thorough in accomplishing every aspect of a task, making sure the models can, like, compose different tools together effectively. Yeah, these are directions we care about. Yeah, I guess it coming back to code. This kind of approach affected the way that we built code also. Because we know that if we want some product that has very broad product market fit today, we would build a cursor or a winter for something like this. These are awesome products that so many people use every day. I use them. That's not the product that we want to build. We want to build something that's much earlier on that curve and something that will maybe be a big product a year from now or have much time from now as the model improves. And that's why code runs in a terminal. It's a lot more bare bones. You have raw access to the model, because we didn't spend time building all this nice UI and scaffolding on top of it. When it comes to the harness, so to speak, and things you want to put around it, there's one that may be prompt optimization. So obviously, I use cursor every day. There's a lot going on in cursor that is beyond my prompt for optimization and whatnot. But I know you recently released compacting contacts, features, and all that. How do you decide how tick it needs to be on top of the CLI? So that's the sound of the share interface. And at what point are you deciding between OK, this should be a part of CLI code versus this is just something for the IDE people to figure out, for example? Yeah, there's kind of three layers at which we can build something. So being a AI company, the most natural way to build anything is to just build it into the model and have the model do the behavior. The next layer is probably scaffolding on top. So as I code itself, and then the layer after that is using Cloud Code as a tool in a broader workflow, so it can compose stuff in. So for example, a lot of people use code with T-mux, for example, to manage a bunch of windows and a bunch of sessions happening in parallel. We don't need to build all of that in. Compact, it's sort of this thing that kind of has to live in the middle because it's something that we want to work when you use code. You shouldn't have to pull in extra tools on top of it. And rewriting memory in this way isn't something the model can do today, so you have to use a tool for it. And so it kind of has to within that middle area. We try to bunch of different options for compacting, rewriting old toolcalls and truncating old messages and not new messages. And then the end, we actually just did the simplest thing, which is ask Cloud to summarize the previous messages, and just return that, and that's it. And it's funny when the model's so good, the simple thing usually works. You don't have to over-engineer it. We do that for Cloud. They spoke them onto just kind of interesting to see that pattern re-emerging. And then you have the Cloud.MD file for the more user-driven memories, so to speak. It's like the equivalent of maybe in cursor rules, I would say. Yeah, and CloudMD, it's another example of this idea of do the simple thing first. We had all these crazy ideas about memory architectures. And there's so much literature about this. There's so many different external products about this. And we wanted to be inspired by all the stuff. But in the end, the thing we did is ship the simplest thing, which is it's a file that has some stuff,

Context, y ahora hay un poco de versiones de desfile, puedes poner en la ruta, o puedes poner en chal directories, o puedes poner en tu casa directoria, y vamos a ver todos estos en diferentes formas, pero sí, simple es que eso podría funcionar. Yo creo que eres familiar con Ada, que es una cosa que hay gente en la escolar de la gente, y luego cuando Cloud Code llegó a la misma gente de Cloud Code. Any thoughts on inspiration that you took from it, things to the differently, maybe design principal, en el que quieras de una forma diferente. Sí, este es, en realidad en el momento de que me ha pagado CGI, esto te arruvía. Ok, voy a decirlo, esto es true. Ada inspired este internal tool que hemos utilizado en Anthropic called the Cloud. So Cloud es como CLI. Y esa es la prenta de la compra de Cloud Code. Es una de las ratas de la escuela de la escuela que es como, esto es como un nuevo Python, que toma un minuto de start up, es como muy mucho rotado por las ratas, no es un producto de publish. Cuando cuando estaba en el first join, leaving out my first bar request. I wrote this bar request, porque no me cilindran muy bien. My bootcamp buddy at the time Adam Wolf was like, you know, actually maybe instead of handwriting, I just asked Bide to write it and I was like, okay, I guess so. A, I love, maybe there's some, you know, capability I didn't know about. And so I start up this terminal tool and it took like a minute to start off and I ask, quite hey, here's the description can you make a PR for me and after a few minutes of chucking along it made a PR and it worked. And I was just blown away. porque no hubo idea, no hubo idea que eran los estudios que podría hacer este tipo de cosas. Así que pensé que, tipo, single line autocompleo, era el estado de Arb, antes de ir adentro. Y entonces, que era el momento en que me dejó a AGIFILD. Y sí, esa es la cosa que fue de Cod. Sí, era ADER, Inspired, Clyde, que Inspired, Cod Code. Es muy grande, fan de ADER, es un producto awesome. Creo que, como personas son interesantes en comparar y contrastar, obviamente, Porque a ti, obviamente, este es la casa de tuyo que trabajas en esto. people are interested in figuring out how to choose between tools. There's the curses of the world. There's like devins of the world. There's haters and there's cloud code. And we can't try everything all at once. My question would be, where do you place it in the universe of options? Well, you can ask cloud to just trial these tools. And I wonder what it would say. No self-favoring at all. Cloud plays engineering. I don't know. We use all these tools in-house, too. We're big fans of all this stuff. Cloud code is, obviously, it's a little different than some of these other tools in that it's a lot more raw. Like I said, there isn't this kind of big, beautiful UI on top of it. It's raw access to the model. It says raw as I get. So if you want to use a power tool that lets you access the model directly and use cloud for automating big workloads, for example, if you have 1,000 wind violations and you want to start 1,000 instances of cloud and have it fix each one, and then make a peer, then cloud code is a pretty good tool. It's a tool for power workloads for power users. And I think that's kind of where it fits. It's the idea of parallel versus kind of single-pad one way to think about it. Where the IDE is really focused on what you want to do versus clock code. You kind of more see it as less supervision required. You can kind of spin up a lot of them. It's at the right mental model. Yeah, and there's some people at Anthropic that have been racking up thousands of dollars a day with this kind of automation. Most people don't do anything like that, but you told we could do something like that. Yeah, we think of it as a unix utility, right? So it's the same way that you would compose GREP or CAT or oh, CAT, just kind of. For some reason, this is the same way you can compose code into workflows. The cost thing is interesting. Do people pay internally? Or do you get free? If you work at Anthropic, you can just run this thing as much as you want every day. It's for free internally. Nice. Yeah, I think if everybody had it for free, it would be huge. What? Because, I mean, if I think about, I pay a cursor 20 bucks a month, I use millions and millions of token and cursor that will cost me a lot more in clock code. And so I think like a lot of people that I've talked to, they don't actually understand how much it costs to do these things and they'll do a task and they're like, oh, that costs 20 cents. I can't believe I've paid that much. How do you think, going back to the product site too, is how much do you think of that being your responsibility to try and make it more efficient versus that's not really what we're trying to do with the tool? We really see clock code as the tool that gives you the smartest abilities out of the model. We do care about cost in so far as it's very correlated with latency. We want to make sure that this tool is extremely snappy to use and extremely thorough in its work. We want to be very intentional about all the tokens that it produces. I think we can do more to like communicate the cost with users. Currently, we're seeing costs around like $6 per day, per active user. And so it's like, it does come out to a bit higher over the course of a month and cursor, but I don't think it's like out of band. And that's roughly how we're thinking about it. I would add that I think the way I think about it is it's our Y question. It's not a cost question. And so if you think about an average engineer salary and we were talking about this before the podcast, engineers are very expensive. And if you can make an engineer 50%, 70% more productive, that's worth a lot. And I think that's the way to think about it. So if you're targeting clock to be the most powerful end of the spectrum, as opposed to the less powerful, but faster, cheaper side of the spectrum, then there's typically people who recommend a waterfall, right? You try this faster, simple one. That doesn't work. You upgrade, you upgrade, you upgrade, and finally you hit clock code. At least for people who are token constrained that don't work it in topic. And part of me wants to just fast track all that. I just want to fan out to everything all at once. And once I'm not satisfied with the next one solution, I'll just switch to the next idea. I don't know if that's real. Yeah, we're definitely trying to make it a little easier to make clock code kind of the tool that you use for all the different workloads. For example, we've launched thinking recently. So for any kind of planning workload where you might have used other tools before you can just ask quad and that'll use a chain of thought to think stuff out. I think we'll get there. Maybe we'll do it this way. How about we recap sort of the brief history of clock code? Between when you launch it now, there have been quite a few ships. How would you highlight the major ones? And then we'll get to the thinking tool. And I think I'd have to like check your Twitter to remember everything. I think a big one that we've gone a lot of requests for is web fetch. So we worked really closely with our legal team to make sure that we shipped as secure of an implementation as possible. So we'll web fetch if a user directly provides a URL, whether that's in their call.md. Or in their message directly. Or if a URL is mentioned in one of the previously fetched URLs. And so this way enterprises can feel pretty secure about letting their developers continue to use it. We shipped a bunch of auto features like auto-complete where you can press tab to complete a file name or file path, auto-compact, so that users feel like they have infinite context since we'll compact behind the scenes. And we also shipped auto-except because we noticed that a lot of users were like, hey, like Cloud Code can figure it out. I've liked it after a lot of trust for Cloud Code. I wanted to just like autonomously edit my files, run tests, and then come back to me later. So those are some of the big ones. VIN mode, custom slash commands. People love VIN mode. So that was a top request too. That one went pretty viral. Yeah. Memory, those are recent ones. Like the hashtag to remember. So yeah, I mean, I'd love to dive into on the technical side any of them that was particularly challenging. Paul from Aitor always says how much of it was coded by Aitor. So then the question is, how much of it was coded by Cloud Code? Obviously there's some percentage, but I wonder if you have a number, like 50, 80? It's very high. Probably near 80. Yeah, it's very high. It's a lot of human code review though. Yeah, a lot of human code review. I think some of the stuff has to be handwritten. And some of the code can be written by Cloud. And there's sort of a wisdom in knowing which one to pick and what percent for each kind of task. So usually where we start is Cloud writes the code. And then if it's not good, then maybe a human will dive in. There's also some stuff where I actually prefer to do it by hand. So it's like intricate data model refactoring or something. I won't leave it to Cloud because I have really strong opinions. And it's easier to just do it and experiment and it is just explain it to Cloud. So yeah, I think that nuts out to maybe 80, 90% quadridden code overall. Yeah, we're hearing a lot of that in our portfolio of companies, like more like Series A companies is like 80, 85% of the code they write is ad-generated. Yeah, yeah. Yeah, well, that's a whole different discussion. The custom slash command I had a question. How do you think about custom slash command, MCPs, how does this all tie together? Is the slash command and clock code kind of an extension of the MCP? Are people building things that should not be MCP, but are just kind of like self-contained things in there? How should people think about it? Yeah, I mean, obviously we're big fans of MCP. You can use MCP to do a lot of different things. You can use it for custom tools and custom commands and all the stuff. But at the same time, you shouldn't have to use it. So if you just want something really simple and local, you just want some essentially like prompt that's been saved. Just use what we'll command for that. Over time, something that we've been thinking a lot about is.

¿cómo re-expose cosas en convenient ways? Por ejemplo, ¿could you re-expose that as anmcp prompt? ¿Por qué Cloud Code es anmcp find and anmcp server? O summary, ¿qué pasa en el custom bash tool? ¿Y es la forma de re-expose that as anmcp tool? Sí. Y creo que en general, no hay que ser una tecnología en particular. ¿Eres algo que funciona? Sí, porque hay algunos papas aquí. Creo que es una buena forma de usar el Cloud Code para testar. Hay un propósito en el mcp protocolo, pero entonces, personas pueden también re-expose las comienzas. Y me gusta ver que el mcp va a ser, donde es casi cada comienzo, lejos en el mcp, pero no es una mcp en el mcp, porque es en un custom bash. Creo que eso es lo que siempre hay que ver. Es como, ¿could this be in the runtime o en el mcp server? Creo que hay personas que hay en que ver el que hay que ver. Sí. Para algunos papas, creo que eso probablemente es en el mcp, porque hay un poco de los tool calls que hay en el mcp. Y así es muy bien que se mueve en el mcp server. O sea, las comienzas son realmente propias, no son actuales. Estamos pensando en cómo explosar más las comienzas y las comienzas, para que las personas puedan hacer sus tus toolas o tornar algunos de tus toolas que el Cloud Code viene con, pero también hay un tricuncio ahí, porque lo que quiero hacer es hacer las comienzas que las personas puedan hacer, o las comienzas que el Cloud es lo que se puede entender, y que las personas no están accidentes, inhibirle su experiencia, por dar un tool que es confusible para el Cloud. Estas están intentando trabajar por el UX. Yo te voy a dar un ejemplo de cómo esto se conecta. Para el Cloud Code en turno, en el GitHub repo, tenemos el GitHub Action que viene, y el GitHub Action en Vox Cloud Code con el command slash local. Y el command slash es Lynth. Así que esto se encuentra en Vox Cloud. Y es un montón de cosas que están muy triki con un tricuncio con un tricuncio tradicional que está en una canalización estaticante. Entonces, por ejemplo, esto es un check para los textuales, pero también es un check de los codes, los comienzas. Esto es un check de los que usamos un library particular para los efectos de network, instead de los built-in library. Hay un montón de estos especificos que vemos, que están muy difíciles de expresar un link. Y en este sitio, puedes ir y ver un montón de link rules para esto. Tienes que ver un montón de link. Pero, honestamente, es mucho más fácil de dar un link en Markdown en un local command. Y simplemente commit eso. Y así que lo que hacemos es que el Cloud viene por el GitHub Action. Nos invocamos con el project slash CullenLint, así que es en Vox, que es el command local. Se encuentra el link. Se identifica en los estéis. Se encuentra en los codes, y entonces se usa el GitHub MCP Server en lugar de que commit le changes a la pieza. Y así, puedes comprobar estos todos. Y creo que es la forma de pensar en el codo. Es solo un todo y el ecosistema que comprobará. Nuestro. without being opinionated about any particular piece. Es interesting. I have a weird chapter in my CV that makes me I was the CLI maintainer for a nullify. And so I have a little bit of a dive. There's a decomplación of Cloud Code out there that has since been taken down. Seems like he used Commander JS in React Ink is the public info about this. Just kind of curious, at some point you're just not even building Cloud Code. You're kind of just building a general purpose CLI framework that anyone, any developer can hack to their purposes. You ever think about this? This level of configurability is more of a CLI framework or some new form factor that doesn't exist before. Yeah, it's definitely been fun to hack on a really awesome CLI because there's not that many of them. But yeah, we're big fans of Ink. Vadim Dmites. So we actually used him, used reacting for a lot of our projects. Oh, cool. Yeah. Yeah, Ink is amazing. It's like, it's sort of hacky and chinky in a lot of ways. It's like you have React. And then the renderer is just translating the react code to like ANC escape codes as the way to render. And there's all sorts of stuff that just doesn't work at all because ANC escape codes are like this thing that started to be written like the 1970s. And there's no really great spec about it. Every terminal is a little different. So building in this way, it feels to me a little bit like a building for the browser back in the day where you have to think about like, Unitex 4 or 6 versus Oprah versus like Firefox and whatever, like you have to think about these cross terminal differences a lot. But yeah, big fans of Ink because it helps abstract over that. We're also, we use bun. So big fans of bun. That's been, it makes writing our tests and running tests much faster. We don't use it in the runtime yet. It's not just for speed, but you tell me I don't want to put words to your mouth. But my impression is they help you ship the compilation, the executable. Yeah, exactly. So we use bun to compile the code together. Yeah. Any other pluses of bun. I just want to track bun versus Deno conversations. Yeah. These Denos in there, you know? I actually haven't used Deno back. It's been a while. Yeah. I remember a lot of people say. Yeah. Ryan made it back in the day. And it was like, there are some ideas that I think were very cool in it. But yeah, I just never took off to that same degree. Yeah. Still a lot of cool ideas like being able to NPM just to import from any URL, I think, is pretty amazing. That's the dream of ESM. Yeah, very cool. OK. I was going to ask you one other feature. Then we can get to the thinking tool of AutoExcept. I have this little thing I'm trying to develop thinking around for trust in agents, right? When do you say, all right, go autonomous. When do you pull the developer in? And sometimes you let the model decide. Sometimes you're like, this is a destructive action. Always ask me. And I'm just curious if you have any internal heuristics around when to AutoExcept. And where all this is going? We're spending a lot of time building out the permission system. So Robert, on our team is leading out this work. We think it's really important to give developers the control to say, hey, these are like the allowed permissions. Generally, this includes stuff like, the model's always about to read files or read anything. And then it's up to the user to say, hey, is about to edit files, is about to run tests. These are probably the safest three actions. And then there's a long list of other actions that users can either allow this or deny this based on REGX matches with the action. How can writing of how ever be unsafe if you have version control? I think that's. Yeah, I think there's like a few different probably like aspects of safety to think about. So it could be useful just to break that out a little bit. So for file editing, it's actually less I think about safety, although there is still a safety risk because what might happen is, let's say, the model fetches a URL. And then there's a prompt injection attack in the URL. And then the model writes a malicious code to disk and you don't realize it. Although there is code review as like a separate where there is a protection. But I think generally for file rights, the model might just do the wrong thing. That's the biggest thing. And what we find is that if the model is doing something wrong, it's better to identify that earlier and correct it earlier and then you're gonna have a better time. If you wait for the model to just go down this like totally wrong path and then correct it 10 minutes later, you're gonna have a bad time. So it's better to use you identify failures early. But at the same time, there's some cases where you just wanna let the model go. So for example, if quad code is, you know, it's writing tests for me. I'll just hit shift tab, enter auto accept mode and just let it run the tests and iterate on the tests until they pass. Because I know that's a pretty safe thing to do. And then for some other tools like Bash tool, it's pretty different because quad could run, you know, RM, RF slash. And that would suck, that's not a good thing. So we definitely want people to be in the loop to catch stuff like that. The model is, you know, trained and aligns to not do that. But, you know, these are the non deterministic systems. So like you used to want a human in the loop. Yeah. I think that generally the way that things are trending is kind of less time between human input. Did you see the meter paper? No, the established a Moore's law for time between human input basically. And it's basically doubling every three to seven months is the idea. And the topic is currently doing super well on that benchmark. It's roughly about autonomous for 50 minutes at the 50% of human effort, which is kind of cool. It's highly recommend that. I put cursor in Yolomode all the time and just run it. But it's fine. Vibe coding, right? Like this is called, let's say the spade. And there's a couple things that are interesting when you talked about alignment and the model being trained. So always put it on Docker container. And I have a prefix every command with like the Docker Compost. And yesterday, my Docker server was not started. And I was like, oh, Docker is not running. Let me just run it outside of Docker. And I'm like, whoa, whoa, whoa, whoa. You should start Docker and run it in Docker. And I'll go outside. So that is like a very good example of like, you know, sometimes you think is doing something and then doing something else. And for the review side, it's all I would love to just chat about them more. I think the linter part that you mentioned, I think maybe people skipped it over. It doesn't register the first time, but like going from like rule-based linting to like semantic linting. I think it's like great and super important. And I think a lot of companies are trying to do how to do autonomous PR review, which I've not seen one that I use so far. They're all kind of like mid. So I'm curious how you think about closing the loop or making that better. And figuring out especially like, what are you supposed to review? Because these PRs get pretty big when you buy code, you know, sometimes I'm like, oh, wow. OGTM. You know, I was like, am I really supposed to read all of this? It kind of seems, most of it seems pretty standard. But like, I'm sure there are parts in there that the model would understand.

que están en la distribución, así que esto es lo que estamos viendo. Así que no es un muy open-up en la pregunta, pero creo que hay que tenerlo. Sí, tenemos algunos experimentos donde Clot es en codervio internamente, estamos muy contentos con los resultados, así que no es algo que queremos abrir muy bien. La forma que estamos pensando en que Clot Code es, como decía antes, es primitivo. Así que si quieres usar el codervio de codervio, puedes hacer esto si quieres hacer un codervio como un secreto, un codervio de vulnerabilidad, un codervio. Si quieres hacer esto, si quieres hacer un codervio como un rompido semántico, puedes hacer eso. Y hoping con codervio, eso hace eso. Así que si quieres hacer esto, si quieres hacer esto, es solo si quieres hacer un codervio de codervio. Y puedes hacer un codervio de codervio de codervio, también, porque el codervio es muy grande en la edad de codervio de codervio. Sí, una pregunta es que tenemos una intelectiva de codervio, que es como un codervio de codervio y como utilizamos codervio en las situaciones para automatizar el codervio. Y también, muchos de los colegios de codervio de codervio tienen una intelectiva de codervio de codervio. Así que, por ejemplo, se dice que hay 100,000 testes en mi repo. Algunas de ellos son de update, algunas de ellos son de fueci, y ellos son de codervio de codervio de codervio para ver a todos estos testes en la parte de la edad. OK, ¿cómo se puede obtener alguna de ellos? ¿Cómo se puede deprimir a ellos? ¿Cómo se puede hacer un codervio de codervio de codervio? Así que eso ha sido una muy buena forma que personas no interactivamente con codervio de codervio. ¿Qué es el mejor practicado aquí? Porque cuando es un codervio de codervio, puede ser todo, y no necesariamente reviewing el output de todo. Así que yo estoy muy curioso cómo es diferente en una intelectiva de codervio de codervio. ¿Qué es el mejor problema de los impuestos para los argumentos? Sí, y por favor no hay una utilidad, así que no es un codervio de codervio de codervio de codervio. Y luego, pasen a los propios propios. Y eso es todo. Es un codervio de codervio de codervio. Por generalmente, es el mejor test que nos vemos. Eso es el lugar donde funciona muy bien. Y, bueno, hay que pensar en la permisión de la forma del primer, y cosas como eso. Así que, por ejemplo, un par de diferentes que nos parece y no se fixen los problemas. O, por ejemplo, estamos trabajando en un tema donde usamos un codervio con DASHP para generar el codervio de codervio de codervio. Así que el PR está mirando sobre la historia de la comandía y preguntando, ¿qué es esto? Pero esto es que está en el codervio de codervio de codervio de codervio. Porque nosotros no sabemos que hay gente que ha buscado un codervio de codervio de codervio de codervio de codervio de codervio. Así que generar una intelectiva de codervio es muy bien para los testes que queramos ver. La cosa que siempre recomiendo es que pasen un muy específico de permisión en la command line. Así que puedes hacer las pasen dash a la altitud. Y entonces puedes dar un específico de tu tuyo. Por ejemplo, no solo a BASH, pero por ejemplo, de la altitud de GIT, o GIT-DIF. Entonces, simplemente dar un set de tuos que puedes usar en la altitud. Esa tuya de default de tuya es una filo, crea que los tuyos como BASH y LAS y Memory Tools. Sí, todos los tuyos. Así que tuya de tuya. Sí, tuya de tuya es de todas las tuyas, pero allowed fools just whats you instead of the permission prompt cause you don't have that in the non interactive mode is just kind of pre-accepting. Yeah, yes. And we also definitely recommend that you start small so like test it on one test, make sure that it has reasonable behavior, iterate on your prompt, then scale it up to 10, make sure that it succeeds or if it fails just like anonymize what the patterns of failures are and gradually scale from there. So definitely don't kick off a run to fix like 100,000 tests. Yeah. Yo creo que en este punto, yo quiero que este tagline es en mi head, que básicamente en el entrope, hay codicode, generar codicode, y codicode también, también en el avión, en el avión, en algún momento, en diferentes personas están presentando, no realmente huérnese en eso, pero eso es lo que ha pasado. Sí, tenemos que ser, en el entrope, hay una humanidad de la avión en el avión, y creo que es para ASL, esto es importante, para el general model de la implementación. ¿Qué es ASL? A, S, L, este es la la de losKelly levels, ¡ да, bien, bien, bien, entonces ¿eo qué es eso? Autonomía, losKelly levels. ¡A Autonomía! Es, es, es, es, es, es, es, es, es, es, es, es, es. Sorry, no, no. Yo no ânimos el acronym, es, es, , okay. Sí, hay un poco de otros, pero por supuesto, ¿Y pasa eso? Sí, no, no sé, no sé lo que es que he called internamente. Sí, exáculo, pero es, es que ¿F Chi, Entonces, ¿Casi el model монones a día es más capable? En la Cía Esa O5 es el hyped level, es, ¿Casi, ¿Esta model es capable de atractarte un user si es un que, ...cante exfaltar a la misma, como se convivir a su contínor... ...porque se publicar a su contínor. Es una eleazar, Yucoski. Ah, y es eso. —Eso es una de las primeras… —¿Y el mundo le va a caer? —We are too, right now. —Yeah, we are kinda bordering on three right now. —Yeah. —So I think I, like, three, four, and five... ...y start having to think a lot more carefully about this... ...because hopefully the model was aligned... ...but in case it's not aligned, you need a human in the loop and the real ways. —The... —Pena de la thinga que me iba a pensar... ...que hay, es, ¿no? en los CTOs, esto es todo muy bien para el individuo de developer, pero las personas que están responsables de la costa de todo, las decisiones de la ingeniería, todo esto va a ir a mi developer, como me llamo de 100 developers, any of them could be doing any of this at this point, what do I do to manage this, how does my code review process change, how does my change management change, I don't know. We've talked to a lot of VPs and CTOs about it, they actually tend to be quite excited, because they experiment with the tool, they download it, they ask it a few questions, and like Cloud Code, when it gives some sensible answers, they're really excited, because they're like, oh, I can understand like this nuance in the code base and sometimes they even ship small features with Cloud Code, and I think through that process of like interacting with the tool, they build a lot of trust in it, and a lot of folks actually come to us and they ask us like, how can I roll it out more broadly? And then we'll often like have sessions with like VPs of dev prod and talk about these concerns around how do you make sure people are writing high quality code? I think in general, it's still very much up to the individual developer to hold themselves up to a very high standard for the quality of code that they merge, even if we use Cloud Code to write a lot of our code, it's still up to the individual who merges it to be responsible for like this being well maintained, well documented code that has like reasonable obstructions. And so I think that's something that will continue to happen where Cloud Code isn't its own engineer, that's like committing code by itself. It's still very much up to the ICs to be responsible for the code that's produced. Yeah, I think Cloud Code also makes a lot of this stuff. A lot of quality work becomes a lot easier. So for example, like I have not manually written a unit test in many months. And we have a lot of units. We have a lot of units. And it's because Cloud writes all the tests. And you know, before I felt like a jerk on someone's PR, I'm like, hey, can you write a test? Because you know, they kind of know they want... For code coverage, is this the relevant? Okay. And you know, they kind of know they should probably write a test and it's probably the right thing to do and somewhere in their head, they make that trade-off or they just want to shift faster. And so you always kind of feel like a jerk for asking. But now I always ask because Cloud can just write the test. Right. You know, so there's no human work. You just ask Cloud to do it and I write the... And I think with writing tests becoming easier and with writing-blit rules becoming easier, it's actually much easier to have high quality code than I was before. What are the metrics that you believe in? Like, is it... A lot of people actually don't believe in 100% code coverage. Because sometimes that is kind of optimizing for the wrong thing. Arguably, I don't know. But like, obviously you have a lot of experience in different code quality metrics. But what still makes sense? I think it's very engineering and it's very engineering team dependent, honestly. I wish there's a one-size-fits-all answer. For me, the one solution. For some teams, test coverage is extremely important. For other teams, type coverage is very important. That's what you... If you're working in a very strictly-type language and, you know, for example, avoiding like any's in JavaScript and Python. Yeah. I think so-comact complexity kind of gets a lot of flak, but it's still, honestly, a pretty good metric just because there isn't anything better in terms of ways to measure code quality. Okay. And then productivity's obviously not lines of code. But do you care about measuring productivity? I'm sure you do. Yeah, you know, lines of code honesty isn't terrible. Yeah, God. It has downsides. Yeah, it's... Well, lines of code is terrible for a lot of reasons. Yes. But it's really hard to make anything better. So you can... It's at least terrible. It's at least terrible. It's like lines of code, maybe like number of PRs. How green your GitHub is. Yeah. Yeah. The two that we're really trying to nail down are one decrease in cycle time. So how much faster are your features shipping because they're using these tools? So that might be something like the time between first commit and when your PR is merged. It's very tricky to get right, but one of the ones that we're targeting. The other one that we want to measure more rigorously is like the number of features that you wouldn't have otherwise built. We have a lot of channels where we get customer feedback. And one of the patterns that we've seen with Cloud Code is that sometimes customer support or customer success will like post, hey, like this app has like this bug and then sometimes 10 minutes later, one of the engineers on that team will be like, Cloud Code made a fix for it. And one of the situations when you like ping them and you're like, hey, that was really cool. They were like, yeah, without Cloud Code, I probably wouldn't have done that because it would have been too much of a divergence on what else other it's going to do. It would have just ended up in this long backlog. So this is the kind of stuff that we really want to measure more rigorously. Those are the other AGI build moment for me. There was a really early version of Cloud Code many, many months ago. And this one engineer at Anthropic Jeremy built a bot that looked through a particular feedback channel on Slack. And he hooked it up to code to have code automatically put a PRs with just fixes to all this stuff. And some of this stuff, you know, it can fix every issue, but it fixed a lot of the issues. And I was like 10%, 50%. You know, this was like early on, so I don't know the number. But it was surprisingly high to the point where I became a believer in this kind of workflow. And I wasn't before.

¿No es el PM? ¿No es que es scary, ¿no? ¿Dónde? ¿Dónde es que puedes hacer tantas cosas? Es como, ¿cómo puedes hacer? ¿Dónde es que hay tantas cosas? Creo que es lo que estoy trabajando con las cosas más, es como, es la habilidad de crear, creer, creer, creer. Pero en algún punto, hay que apoyar, apoyar, apoyar, apoyar. Esto es el parque de Jurassic Park. La gente es muy preocupada con lo que puedes hacer. Sí, sí, exactamente, pero no es lo que deberíamos. Sí, ¿cómo puedes hacer decisiones? No es la costa de implementar la cosa. Es que en el PM, ¿cómo puedes hacer lo que es actual? Sí, también hay una gran barba de neta. Muchos de los fixes, como, hey, este funcionamiento se ha cambiado, o es como, hay un poco de edad, que no hemos encontrado, así que es muy bien, muy bien, muy bien, muy bien, como podemos hacer algo completamente neta. Para neta, creo que hay una gran barba de neta que es muy intuitiva para usar, la experiencia nueva es como, minimal, es como, obviamente, que funciona. Porque tampoco se puedenacyar como una madrota de modos, comoroots, el proceso de nosotros, para asegurar que el futuro definitivamente se fit en la productividad. Es interesante cómo, como es fácil de hacer esto, se transforma la forma que me raro el software, donde, como me dijo, antes de que me raro un gran diseño, y yo creo que era un problema durante mucho tiempo, antes de que me dices algo para algunos problemas. Y ahora, yo recuerdo el código de prototipo, tres versiones de esto, y yo intento el futuro y ver lo que me gusta mejor, y entonces, eso informe mucho mejor y mucho más, doesn a knock or have and i think we haven't total a in turn wise that transition yeah in the industry yeah i feel the same the same way for some tools i build internally people ask me could we do this and i'm like i'll just yeah just build it i feel yeah i feel, it feels pretty good. we should like polish it you know or sometimes it's like nao, that's not broadcasting that you know like that you're guns your max costi even at and that i compete directly on the media that cost is roughly $6 a day que nos da gente con un piezo de pie. Porque yo estoy, como, 6 dollars a día? En fin, 600 dollars a día, tenemos que hablar. Ya, yo pongo 200 bucks a un monstruo de mi studio que me llamo a las fotos. Eso es todo bien, eso es totalmente importante. ¡Y tienes una etapa interna. Y eso es una muy grande lucha que estamos viendo en el mar. Porque muchas veces, si estás trabajando en algo operativo intensivo, si puedes tener un interno de la etapa interna o como una etapa interna donde puedes, por ejemplo, gran acceso a 1000 emails a once. A lot of these things, you don't really need to have a super polished design. You kind of just need something that works. And call codes really good at those kinds of zero to one tasks. Like, we use Streamlit internally and there's been a proliferation of how much we're able to visualize and because we're able to visualize it, we're able to see patterns that we wouldn't have otherwise if we were just looking at raw data. Sí, I was working on also this side website last week and I just showed Cloud Code the mock. So I just took the screenshot I had, dragged and dropped it into the terminal and I was like, hey, what, here's the mock. Can you implement it? And it implemented and it sort of worked. It was a little bit crummy. And I was like, all right, now look at it and puppeteer and iterate on it until it looks like the mock. And then it did that three or four times and then the thing looked like the mock. Yeah, this is just all manual work before. I think we're going to ask about like two other features of I guess the overall agent's pieces that we mentioned. So I'm interested in memory as well. So we talked about auto-compact and memory using hashtags and stuff. My impression is that you're, like you say, simplest approach works. But I'm curious if you've seen any other requests that are interesting to you or internal hacks of memory that people have explored that you might want to surface to others. There's a bunch of different approaches to memory. Most of them use external stores of various sorts. There's a lot of projects like that. And yeah, it's either a k-value or kind of like graph stores that's like the two big shapes for this. You believe in knowledge graphs for this stuff or? You know, I'm a big, if you talked to me before I joined Anthropic and this team, I would have said, yeah, definitely. But now actually I feel everything is the model. Like that's something that wins in the end. And it just as the model gets better, it subsumes everything else. So, you know, at some point the model will encode its whole knowledge graph. It'll encode its own like KVStore if you just give it the right tools. Yeah. But yeah, I think the specific tools, there's still a lot of room for experimentation. We just, we don't know yet. In some ways, are we just coping for lack of context length? Like are we doing things for memory now that if we had like 100 million token context window, we don't care about. It's an interesting way to do that. I would also have 100 million. It's happening context for sure. Some people have claimed it to have done it. We don't know if it's true or not, but... I guess here's the question for you. Sean, if you took all the world's knowledge and you put it in your brain, and let's say, you know, there is like some treatment that you can get to make it so your brain can have any amount of context. You have like infinite neurons. Is that something that you would want to do or would you still want to record knowledge externally? Putting it in my head is like different for me trying to use an agent tool to do it because I'm trying to control the agent. And I'm trying to make myself unlimited but I want to make the tools I use limited because then they deny not to control them. And it's not even like a safety argument. It's just more like... I want to know what you know. And if you don't know a thing, sometimes that's good. Like the ability to audit what's your time. And I don't know if this is the small brain thinking because this is not very bit or less it, which is like, actually sometimes you just want to control every part of what goes in there in the context. And the more you just, you know, Jesus say, the real trust the model, then you have no idea what it's paying attention to. Yeah, I don't know. I see the mac interpretability, stuff from Chris Ola and the team that was part of it. Like last week, yeah, actually. Yes, what about it? I wonder if something like this is the future. So there's an easier way to audit the model itself. And so if you want to see like, what is stored, you can just audit the model. Yeah, the main thing is that they know what features activate it per token and they can tune it up, suppress it, whatever. But I don't know if it goes down to the individual, like item of knowledge from context. You know, not yet. But I wonder, you know, maybe that's the bitter Western version of it. Right. Any other comments from memory? The other ways we can move on to planning and thinking. We've been seeing people play around with memory in quite interesting ways, like having called write a logbook of all the actions that it's done. So that over time, called develops this understanding of what your team does, what you do within your team, what your goals are, how you like to approach work. We would lots of figure out what the most generalized version of this is so that we can share broadly. I think with things like code code, I think when we're developing things like code code, it's actually less work to implement the feature and a lot of work to tune these features to make sure that they work well for general audiences like across a broad range of use cases. So there's a lot of interesting stuff with the memory and we just want to make sure that it works well out of the box before we share it broadly. Agreed that. I think there's a lot more to be developed here. I guess a related problem to memory is how do you get stuff into context? Knowledge-based. Right. And originally, we tried very, very early versions of code actually used RAG. So we indexed the code base and I think we're just using voyage so just off the shelf, RAG, and that worked pretty well. And we tried a few different versions of it. There was RAG and then we tried a few different kinds of search tools. And eventually, we landed on just agentic search as the way to do stuff. And there were two big reasons, maybe three big reasons. So one is it helped perform everything. Which is by a lot. By a lot. And this was surprising. In what benchmark? This is just vibes. Internal vibes. There's some internal benchmarks also, but mostly vibes. It just felt better. In agentic RAG, meaning you just let it look up in however many search cycles it needs. Yeah, just using regular code searching in a glob graph, just regular code search. Yeah. So there was one. And then the second one was there was this whole indexing step that you have to do for RAG. And there's a lot of complexity that comes with that because the code drifts out of sync. And then there's security issues because this index has to live somewhere. And then what if that provider gets hacked? And so it's just a lot of liability for a company to do that. Even for our code base, it's very sensitive. So we're kind of, we don't want to upload it to a third party thing. It could be a first party thing, but then we still have this out of sync issue. And agentic search just sidesteps all of that. So essentially at the cost of latency and tokens, you now have really awesome search without security downsides. Well, memory is like planning, right? There's kind of like memories like what I like to do. And then planning is like, now use those memories to come over to plan to do these things. There was one. Or maybe put it as like memory sort of the past, like what we already did. And if planning is going to what we will do, and it crosses over at some point. Yeah. I think the maybe slightly confusing thing from the outside is what you define as thinking. So there's like a extensive thinking. There's the thing tool. And it's kind of like thinking as in planning, which is like thinking before execution. And then there's like thinking why you're doing, which is like the thing tool. Can you maybe just run people through the difference I already confused listening to you. Why? Well, it's one tool. So quad can think if you ask it to think. Generally, the usage pattern that works best.

es que puedes hacer un poco de la de la search, hacer una aplicación, poner una codiceta en la contexto y luego preguntar a pensar. Y luego puedes hacer un plan de plan de plan, y hacer un plan de plan de semana antes de que se exibiera. Hay algunos propios que tienen una moda de explosiva. Como Ruco, Hastis y Cuín, Hastis. Algunas otras propias de plan de plan de plan de acto o en algunos diferentes modos. Estamos pensando en este approach, pero creo que nuestro approach es una moda de artista a nuestro modelo, que es un poco más. Entonces, fíjate de ser muy simple, fíjate casi al metal. Entonces, si quieres clotar, pide de pensar, por ser como, puedes hacer un plan, piensar, no pide de cualquier código, y eso es la generación de que hablo. Y puedes hacer eso también como puedes. Entonces, hay un plan de estado, y luego clotar un cote o algo, y luego puedes preguntar a pensar en plan un poco más. Puedes hacer eso en cualquier momento. Sí, yo estaba cerrado a la TincTool blog post, y me dijo, ¿Cuál es algo similar para acceder a pensar en un concept diferente? Acceder a pensar en lo que es lo que es lo que hace antes de generar. Y luego pensar en que, once de vez de generar, ¿cómo vas a estar? Y pensar, ¿qué es todo lo que ha hecho en el clonco de Harness? Entonces, no hay que pensar en la diferencia de la differencia. ¿Dónde, tú, básicamente, es la idea? Sí, no hay que pensar en eso. OK, eso es todo. Eso es todo, eso es todo. Porque, a veces, no sé quién es, ¿qué es todo? Bueno, sí, es el clonco de Harness. Sí. En el clonco, en el clonco de Harness, no we don't use a pensar en algo. ¿En etapaoons de clonco piensan a fragile el clonco de Harness? Yo han hecho, a veces, una cosa que hemos pasado antes de recordar. Cuales esc recording y encarcalaron cómo le conocemos en hackathon, que ha hecho un estudio que ya comenzamos a hacer más que un enviante. Que haya hecho que tengamos que poner un poco de VPN la declare en el clonco de Aeodea o lo unir para eso... en la plana. Y entonces, lo que yo le decía el día de hoy, fue básicamente que es muy expensive para siempre hacer eso en el tiempo, pero si te da un tool para que creen y prompide en ciertos casos para usar ese tool, parece ser lo que hace. Ahora, yo creo que es muy curioso, es que tu tarea en general, como el sandbox, el environment, el rendimiento, la reunión de la habilidad, es que es un día que yo me ha hecho una vez que no pensaba que era útil para el cloud, o que el cloud no ha pensado en las opiniones. Yo puedo hablar por horas sobre este cuadro que podemos, ¿no? Sí, sí, sí. Vamos a tener original tokens para mí y luego podemos traer un cloud en ahí. Por supuesto, eso es como expresar lo que este podcast es, o generar tokens para mí. ¿Es el pre-training o el pre-training? Sí, el pre-training es eso. ¿Vamos a tener ahí? Oh, man. Sí, ¿cómo hay? ¿Cómo hay? ¿Cómo hay? ¿Cómo hay que tener un token? Starting with sandboxing. Idealmente, la cosa que queremos es que tendremos un código en un container. Y luego, tendremos un freedom. Y puedes tener un snapshot, con otras cosas, donde puedes tener un snapshot, rewind, todo todo esto. Por supuesto, trabajar con un container con todo esto es como un poco de trabajo, y lo más importante es hacer esto. Y así, queremos, de alguna manera, similizar algunas de estas cosas sin tener una continuación. Hay algunas cosas que puedes hacer hoy. Así que por ejemplo, algo que hago sometimes es, si tengo un cuestión de plan o un cuestión de research, me voy a dejar un cloud de investigar un poco de pasos y un paro. Y tú puedes hacer esto hoy, si te vas a preguntar. Así que, ¿cómo hay que hacer esto? ¿Cómo puedes hacer tres ideas para cómo hacer esto? ¿Cómo hay que hacer esto en parallel? Hay tres agentes para hacer esto. Y en el UI, cuando puedes ver una tasa, es como un cloud de sub. Es un cloud de sub-agentura que hace esto. Y cuando hago algo en Harry, me voy a investigar 3 veces o 5 veces o cada vez en parallel. Y luego el cloud va a poner la mejor opción y se summariza eso para ti. Pero ¿cómo hay la mejor opción de la mejor opción? ¿Qué quieres? ¿Qué se está haciendo de las restas con la etapa internationally? Ahí es la etapa. Pues la etapa de Sapphire está ahí, entonces ya ves, no si, no 같아요. Creo que para nosotros lo estoy handing el hangar. Entonces voy a regular ese modelo. Activ évidemment. carefully de la etapa. Cuando yo se los digo como como decir una etapa y luego mis diferenci additions chicos ya tenís que hacer proteger . Heito prostitucional y sentiamente que exterten las www.codcoedfailing.com es un model muy persistente. Es muy, muy motivado a comprobar el escuelo. Pero es que Sometimes takes the user's goal muy literalmente. Y así, no siempre fulfilling lo que la impuede. Partido de la requesta, porque es así de narro de en, como, yo tengo que... ¡Es lo que lo he hecho! ¡Vamos a hacer! Y así, vamos a probar, bueno, ¿cómo nos da un poco más common sense para que nos den la linea de que estamos tratando de ser muy difícil y no la user definitivamente no quiere eso. Sí, como el clásico example es que, ¡Vamos a probar este testo! Entonces, como, 5 minutos later, es como, bueno, me encarco todo el testo. No, no es lo que me encarco todo el tensor. Pero eso es algo, como, esto es lo mejor de aquí. Como, estos escuelos trabajan sometimes, no, no cada vez, y, como, el model Sometimes tries es muy difícil, pero esto es lo mejor. Sí, como, el contexto, por ejemplo, es un gran uno, donde, a veces, si hay una conversación muy bien, y compácte varios veces, quizás, un de tus intentes originales no era así, como, por supuesto, como cuando lo primero empezaron. Y así, también, el model, como, ¿qué se lo contó? ¿Qué tu original te contó? Y así, estamos muy contentos sobre cosas como, grandes, efectivas, contextos, así que, puedes tener, como, los cargos, como, realmente, tantos de 1000 de tonos, son los testos de cada vez. Y asegúrate de que la codicota se atraen la otra forma. Así que eso sería un profundo, no solo para el Cod Code, pero para todo el Coding Company. Es una historia de David y él es Quinoa de este día. Él es el común de la sensación de 3.5, porque 3.7 es muy persistente. 3.5 se ha tenido algunas historias intercambiadas que, apparently, se ha dejado en tasas y 3.7 no. Y cuando Cod 3.5 se ha dejado, se ha dejado en la primera forma para los trabajadores de el juego, ¿Fix the game? Y ya son screenshots de eso, ¿sabéis? Entonces, si te escuches a él, puedes encontrarlo en YouTube, porque vamos a postar. Muy, muy cool. Una forma o fiel, en el que me da una captura, fue algo que te menciones lo que estamos haciendo en el café, que es que Cloud Code no tiene that much between session memory, o caching, o algo que te llamas, ¿sabéis? Así que, reforma la whole-state por el café de un tiempo. Así que tiene que hacer los mínimos de los cambios que pueden haber en el que hay. ¿Cómo consistencia es el estado, ¿sabéis? Creo que una de las fieles es que lo que se lo hace en el pasado, antes de que te explica la optancia de la CLAW.MD, ¿sabéis de algo que te preocupes? Esa es una cosa que estamos trabajando. Creo que como es el mejor advise ahora, para los que quieren resumen en la CLAW, es decir, CLAW, a ver cómo está la estación de la CLAW, en este texto, Probably not the calls on MD, but in a different docs. And in your new session, tell Claw to read from that doc. But we plan a build in more native ways to handle this specific workflow. There's a lot of different cases of this, right? Sometimes you don't want Claw to have the context. And it's sort of Git. Sometimes I just want a fresh branch that doesn't have any history. But sometimes I've been working on a PR for a while and, like, I need all that historical context. So we kind of want to support all these cases. es un poco más de un poco más, pero generalmente nuestro approach para poder hacer la 다양한idad de la gente por favor con la configuración extra. Así que, once tenemos que hacer algo. ¿Ustedes deberían hacer una tarmo de paro de enlímpos? ¿Cómo vamos? ¿Cómo estamos aquí? ¿Cómo están? ¿Cómo están a la historia? ¿Cómo están a la pantalla de la ropa de la ropa? ¿Cómo están a la ropa de la ropa de la ropa? ¿Cómo están a la paro de la ropa? ¿Cómo están a la ropa de la ropa? para algunas cosas que se han visto en la historia. Por ejemplo, si se dice que si te aplauden, ¡hey, me puse el PR para mí! Se va a ver todo el cambio desde que el brancho se diverge por la mano. Y entonces, te va a ver todos los que no se vean en el primer mensaje. ¿Puedes saber cómo se va a hacer el dip como se va usando? Creo que es muy bueno de traer el cambio que ha pasado en el brancho y así que es, como, lo entiendo antes de continuar con la tasca. Una cosa que otras personas han hecho es hacer un cuadro de commit después de todo el cambio. Y puedes poner esto en el cuadro de Md. Pero hay algunas cosas, como Power User Workforce, que creo que es súper interesante. Estas personas están creando un cuadro de commit después de todo el cambio para que se puedan verlo muy fácilmente. Otas personas están creando un cuadro de crear un cuadro de un cuadro de cada vez, así que se pueden hacer un cuadro de curvas en Paraguay, en el mismo repo. Yo creo que por el punto de vista, queremos apoyar todo esto. Así que, el cuadro de curvas es como un permitido, y no es lo que su trabajo es, eso podría ser. Yo creo que el 3.5, Haiku era el número 4 modelo en Eiter, cuando llegué. ¿Quién es el cuadro de curvas en un mundo en el que haces un cuadro de commit, que haces un cuadro de curvas, para hacer un cuadro de curvas, y cosas como eso, continuamente, y luego hay el 3.7. Eso es lo más. Sí, tú puedes hacer esto si quieres. Así que tú crees un cuadro de curvas, o un cuadro de curvas, o un cuadro de curvas, o un cuadro de curvas? Sí, sí, sí. Bueno, como como un cuadro de curvas, como un cuadro de curvas en el que haces, yo quiero hacer todo esto localmente, como antes que el cuadro de curvas. Sí, así que puedes hacer esto hoy si quieres, así que en el...

you know, if you're using like husky or like whatever precommitted system you're using, or just like get precommitted hooks, just add a line, claw dash P, and then, you know, whatever instruction you have that'll run everytime. I mean, just specify a hi-go. It's really no difference, right? It's like maybe it'll work a little worse, but like it's still supported. Yeah, you can override the model if you want. Generally, we use on it. We default this on it for almost everything, just because we find that it outperforms. Yeah. But yeah, you can override the model if you want. Yeah, I don't have that much money to run. Comethux. Comethux. Controventa. Just like this. As a sign on precomit hooks, I have worked in places where they insisted on having precomit hooks. I have worked in places where they insisted on never do precomit hooks because they get in the way of committing and moving quickly. I'm just kind of curious. Like, do you have a stance or recommendation? Oh, God. That's like asking about tabs versus spaces with a little bit. But like, you know, I think it is easier in some ways to like if you have a breaking test, go fix the test. With clock code. In other ways, it's more expensive to run this at every point, so like these trade-offs. I think for me, the biggest trade-off is you want the precomit hook to run pretty quickly so that if you're either if you're human or if you're a quad, you don't have to wait like a minute for all the types around. So you want the fast version. So generally, you know, precomit, you know, for our code base, should run. Yeah, it's like less than, you know, five seconds or so. Like, just types and went maybe. And then more expensive stuff. You can put in the GitHub action or GitLab or whatever you're using. Good. I don't know. I like putting prescriptive recommendations out there so that people can take this and go like, this guy said it, we should do it in our team. And that's a basis for decisions. Yeah, yeah, yeah. Cool. Any other technical stories to tell? You know, one that's zoomed out into more productsy stuff, but you can get as technical as you want. I don't know. Like, one anecdote that might be interesting is the night before the code launch, we were going through to burn down the last few issues and the team was up like pretty way trying to do this. And one thing that was bugging me for a while was we had this like markdown rendering that we were using. And it was just, you know, it's like the markdown rendering in quad today is beautiful. And it's just like really nice rendering in the terminal when it is bold and, you know, headings and spacing and stuff very nicely. But we tried a bunch of these off-the-shelf libraries to do it. And I think we tried like two or three or four different libraries. And just nothing was quite perfect. Like, sometimes the spacing was a little bit off between a paragraph and like a list. Or sometimes the text wrapping wasn't quite correct. Or sometimes the covers weren't perfect. So each one had all these issues. And all these markdown renders are very popular. And they have, you know, thousands of stars on GitHub and have been maintained for many years. But, you know, they're not really built for terminal. And so the night before the release, I like 10 PM on my mic, all right, I'm going to do this. So I just asked quad to write a markdown per se for me. And they wrote it. Zero shot. Yeah. It wasn't quite zero shot, but after, you know, like maybe like one or two prompts, they got it. And, you know, that's the markdown parser that's encoded in. The reason that markdown books were beautiful. That's a fun one. It's interesting what the new bar is. I guess for implementing features, like this exact example where there's libraries out there that you normally reach for that you find, you know, some dissatisfaction with for literally whatever reason. You could just spin up an alternative and go off of that. Yeah. I feel like AI has changed so much and, you know, a lot of these problems are, you know, like the example we had before a feature you might not have built before or you might have used a library and now you can just do it yourself. Like the cost of writing code is going down and productivity is going up. And we just have not internalized what that really means yet. Yeah. But yeah, I expect that a lot more people are going to start doing things like this, like writing your own libraries or just shipping every feature. Just to zoom out, you obviously do not have a separate cloud code subscription. I'm curious what the roadmap is. Like this is just going to be a research for view for much longer. Are you going to turn it into an actual product? I know you were talking to a lot of CTOs and BPs. Is there going to be cloud code enterprise? What's the vision? Yeah. So we have a permanent team on cloud code. We're growing the team. We're really excited to support cloud code in the long run. And so, yeah. Well, we plan to be around for a while. In terms of subscription itself, it's something that we've talked about. It depends a lot on whether or not most users would prefer that over pay as you go. So far, pay as you go has made it really easy for people to start experiencing the product because there's no upfront commitment. And it also makes a lot more sense with a more autonomous world in which people are scripting cloud code a lot more. But we also hear the concern around, hey, I want more price predictability if this is going to be my go-to tool. So we're very much still in the stages of figuring that out. I think for enterprises, given that cloud code is very much like a productivity multiplier for ICs. And most ICs can adopt it directly. We've been just supporting enterprises as they have questions around security and productivity monitoring. And so, yeah, we've found that a lot of folks see the announcement and they want to learn more. And so we've been just engaging in those. Do you have a credible number for the productivity improvement? Like, for people who nodded in topic that you've talked to, like, you know, are we talking 30%? Some number would help justify things. We're working on getting this. We should, yeah. We're actually working on it. But anecdotally, for me, it's probably 2X, my productivity. Oh my god. So I'm just like, I'm an engineer that codes all day every day. For me, it's probably 2X. Yeah. For some engineers at Anthropic, we're probably 10X, their productivity. And then there's some people that haven't really figured out how to use it yet. And, you know, they just use it to generate, like, commit messages or something. That's maybe, like, 10%. So I think there's probably a big range and I think we need to study more. For reference, sometimes we're in meetings together and sales or compliance or so on. It's like, hey, like, we really need, like, X feature. And then for us, we'll ask a few questions to, like, understand the specs. And then, like, 10 minutes later, he's like, all right. Well, it's built. I'm going to merge it later. Anything else? No. So it definitely feels definitely far different than any other PM role I've had. Do you see yourself opening that channel of the non-technical people talking to Cloud Code and then the instance coming to you, which like, they're ready to find and talk to it and explain what they want. And then you're doing kind of the review side and implementation. Yeah. We've actually done a fair bit of that. Like, Megan, the designer on her team, she is not a coder, but she's winning poor requests. She uses code to do it. She designs the UI? Yeah. And she's landing PRs to our console product. So it's not even just like building on Cloud Code. It's building like across our products suite in our Mono repo. Right. Yeah. And similarly, you know, our data scientists uses Quad Code, right? Like, you know, like BigQuery queries and there was like some finance person that went up to me the other day and was like, hey, I've been using Quad Code and I'm like, what? Like, how do you get installed? You know, you can use Git. And they're like, yeah, yeah, I figured it out. And yeah, they're using it. They're like, so Quad Code, you can pipe in because it's a unique security. And so what they do is they take their data, put it in a CSV, and then they take the cat, the CSV, pipe it into code. And then they ask it code questions about the CSV and they've been using it for that. Yeah. That would be really useful to me. Because really what I do a lot of the times, like somebody gets me a feature request, I kind of like rewrite the prompt. I put it in agent mode and then I review the code. It would be great to have the PR wait for me. I'm kind of useless in the first step. Like, you know, taking the feature request and prompting the agent to write it, I'm not really doing anything. Like, my work really starts after the first run. It's done. I was going to say, like, I can see it both ways. So, like, okay, so maybe I'll simplify this to, in the workflow of non-technical people in loop, should the technical person come in at the start? Or come in at the end, right? Or come in at the end and the start, obviously, that's the highest leverage thing. Because like, sometimes you just need the technical person to ask the right question. That the non-technical person wouldn't know to ask. And that really affects the implementation. But isn't that the better lesson of the model, that the model will also be good at asking the follow-up question? Like, you know, if you're like telling the model, hey. That's what it says. You trust the model to do the least, right? Yeah. So, go ahead. Yeah. The model, hey, you are the person that needs to translate this non-technical person request into the best prompt for cloud code to do the first implementation. Yeah. Like, I don't know how good the model will be today. I don't have any value for that. But that seems like a promising direction for me. Like, it's easier for me to review 10 PRs than it is for me to take 10 requests, then run the agent 10 times, and then wait for all of those runs to be done and review. I think the reality is somewhere in between. We spend a lot of time shadowing users and watching people at different levels of seniority and technical depth use code. And one thing we find is that people that are really good at prompting models from whatever context, maybe they're not even technical, but they're just really good at prompting. They're really effective at using code. And if you're not very good at prompting, then code tends to go off the rails more and do the wrong thing. So I think in this stage of where models are at today, it's definitely worth taking the time to learn how to prompt models well. But I also agree that maybe in a month or two or three months, you won't need this anymore because the better lesson always wins. Please. Please do it. Please do it on traffic. I think there's a broad interest in people forking or customizing cloud codes. We have to ask, why is it not open-source? We are investigating. So it's not yet. There's a lot of trade-offs that go into it. On one side, our team is really small and we're really excited.

para la contribución de open-source, si es open-source, pero es un poco de trabajo para mantener todo, y verlo, verlo y mantener un poco de open-source, y un poco de otras personas en el team-do-2. Y es un poco de trabajo, como es un tiempo de trabajo, en la contribución, y todo esto. Sí, yo me voy a explicar que puedes hacer un trabajo available. Y eso es, ¿sabes un poco de individuos de uso, sin ir a la parte legal de todo el juego de open-source? Sí, exactamente. Yo creo que creo que hay algo de secreto en el esfuerzo. Y, obviamente, es todo el JavaScript, que puedes compárelo. Compárelo. Es muy interesante. Sí. Y en general, todo el secreto de todo el juego de open-source es todo en el model. Y este es lo que es posible para el model. Nosotros no podemos dar nada más minimal. Este es lo más minimal. Sí. Entonces, no es mucho. Si hay una arquitectura que podrías estar interesada en, eso es no simple. ¿Qué piste? Es una alternativa. Estamos hablando de una arquitectura de gente agente aquí, ¿sabes? Hay un loop aquí, y él va a poner en los modelos y en una relativa relativamente intuitiva. Si vienes a ir a la descripción de la descripción y choque la generación de la parte legal, ¿qué es lo que es? Pues Boris ha revertido en este. Boris en el equipo ha revertido en este tipo de 5 veces. Oh. Eso es una historia. Sí, es muy importante. Es muy importante. Es la más simple. Creo que por la design. OK, esto es simple. Sí, tiene que ser más complejo. Pero si hemos retenido en la descripción, probablemente cada 3 o 4 semanas o algo. Y es como todo es una serie de fiestas. Todo el tiempo se queda siendo suave. Y es porque la clara es tan bonita en la descripción. Sí, la verdad es que la forma de la transformación es la interfaz. La clara de la MCP. Pero todo eso tiene que estar en la forma de la transformación hasta que realmente hay una gran razón de la transformación. Sí. Creo que las las transformaciones son más simples, como hacer interfaces con diferentes componentes. Porque, últimamente, queremos hacer que el contexto que es hecho en el model es en la pura forma y que la harnessa no se interviene con la intenta user. Y así, muy bien, un poco de eso es como remover cosas que podría ser la forma o que podría confiar el model. Sí, en el UX hay algo que ha sido muy triste y la razón de que, ya sabes, tenemos un designer trabajando en el terminal app. Es realmente muy difícil� hacer un terminal. Es como no hay un literature. Como ha sido haciendo un producto durante hace hace años. Tienes una idea de cómo hacer apps y web, para ingenieros, en términos de como se ha que hay devex. Pero el terminal es una razón. Tienes un gran sistema de terminal, como usado de cursos y cosas, para físico y y sistemas de UI. Pero estos son todo el que se están muy antiquados por el UIs-tendredo de hoy. Y así,利le mucho el trabajo de hacer cómo exactamente hacer el ápio, como frech y moderno y intuitivo en el terminal. Y tenemos que ir con la idea de diseñar el lenguaje. Sí, yo creo que te voy a desarrollar por el tiempo. Es una pregunta clases. Es más general, creo que a veces hay gente que está listo. Y en tal vez, creo que es fácil de decir que es el mejor brand para AI-Engineering, como desarrolladores y code models. Y ahora, con el code tool, en actualmente, es, es, es, es el pelo de todo, model, tool y protocolo. ¿Como? Y no sé, esto es ovea, un año ago, en el cual, 3 de la vez, era más, como, este es el general propio, no es eso, pero el clon en el real, en el cual, era el tipo de cómo de tuelos. Y creo que ha estado en el propio espleno, y ustedes no están extendiendo. Entonces, ¿Y qué haces el parque? Eso es que es el mejor que haces. Esto parece como que no se centralizan, En el tiempo, cuando hablaban a personas, They were like, oh, yeah, we just had this idea. We pushed it and it did well. Y no, I'm just, There's no centralized strategy here. Or, is there an overarching strategy? Sounds like a PM question. I don't know. I would say, Dari was not breathing down your necks going like, build the best deaf tools. He's just letting you do your thing. Everyone just wants to do awesome stuff. It's like, I feel like the model just wants to write code. Yeah, I think a lot of this trick goes down from the model itself, being very good at code generation. We're very much building off the backs of an incredible model. That's the only reason why code code is possible. I think there's a lot of answers to why the model itself is good at code, but I think one high-level thing would be so much of the world is run via software. And there's immense demand for great software engineers. And it's also something that you can do almost entirely with just a laptop or just a deaf box or some hardware. And so it just is an environment that's very suitable for LLMs. It's an area where we feel like you can unlock a lot of economic value by being very good at it. There's a very direct ROI there. We do care a lot about other areas, too, but I think this is just one in which the models tend to be quite good. And the team's really excited to build products on top of it. And you're growing the team you mentioned. Who do you want to hire? Yeah, we are. Who's like a good fit for your team? We don't have a particular profile. So if you feel really passionate about coding and about the space, if you're interested in learning how models work and how terminals work and how all these technologies that are involved, yeah, hit us up. We'll be happy to chat. Awesome. Well, thank you for coming on. This was fun. Thank you. Thanks for having us. This was fun.