CL Kao (Recce) on What Breaks When AI Touches Your Data - Episode 6
Your agent wrote the migration and the diff looks fine. What did it do to the millions of rows it touched, and to every dashboard, finance report and model that reads them downstream? Today's guest has spent years on what correctness means when the thing changing is data, and lately, when the thing making the change is an agent.
CL Kao built SVK, one of the earliest decentralized version control systems, about two years before git. He co-founded g0v, Taiwan's civic tech community, and competed for Taiwan at the International Olympiad in Informatics. Today he's the founder and CEO of Recce, a data review tool for dbt pipelines backed by Heavybit, and the creator of Spacedock, a workflow toolkit for agents with approval gates and review built in.
They start with data review: what a reviewer actually sees, why diffing two versions of your data gets noisy fast, and the blast radius view that shows every downstream column a change touches. Then they get into Spacedock, which came out of a very specific kind of exhaustion: running five agents at once and becoming the human connector between all of them.
CL also walks through lure fixtures, honeypots in Spacedock that catch an agent over-building. His favorite example is an agent that decided to write a whole terminal emulator to test a markdown reader when tmux was sitting right there. He also explains Behavior Diff, a lightweight way to see how a change to a skill or an AGENTS.md shifts what your agent does before you build a full eval suite.
Plus: the agent that refused to rebase, why every IC is a manager now, borrowing sprint planning for agent teams, and the sandbox he runs his own agents in.
Sites:
Recce ➡️ reccehq.com
Spacedock ➡️ spacedock.md
Agent Safehouse ➡️ agent-safehouse.dev
CL on LinkedIn ➡️ linkedin.com/in/clkao
Taskless ➡️ taskless.io
Topics covered:
- From competitive programming to SVK, two years before git
- What a data review is, and what it catches that code review misses
- What a reviewer actually sees: profiles, value diffs and column-level lineage
- Migrations versus intended changes, and why the second kind is harder to verify
- What breaks when agents touch your data
- The customer who told Recce "you don't have to sell me this"
- Blast radius: every downstream column a change touches
- Custom rules for how much evidence each model needs
- Spacedock and the exhaustion of running 10 agents at once
- Deciding when a human should get involved
- Lure fixtures and the agent that built a terminal emulator
- Behavior Diff: seeing how a skill change shifts agent behavior
- Sprints for agent teams, without the capacity cap
- Two times AI surprised him, and sandboxing with Agent Safehouse
Timestamps:
00:00 Welcome to The Task at Hand
00:32 CL's background: SVK to data
03:34 What is data review?
07:20 What reviewers actually see
08:56 When AI touches data
11:02 What breaks with agents
12:11 Recce customer stories
14:07 Blast radius visualization
16:32 Custom rules and risk
18:11 Spacedock: agent workflows
20:22 Agent exhaustion and context
25:35 Managing your attention
29:06 Spacedock at Recce
30:50 Lure fixtures and over-engineering
37:27 Productionizing agent ideas
40:30 AI surprises and sandboxing
Transcript
[00:00:00] Welcome & Introduction
[00:00:00] Jakob Heuser: Welcome everybody to The Task at Hand. I'm Jakob Heuser, CTO co-founder at Taskless. Today, I have with me CL Kao, creator of Recce and Spacedock, talking about the idea of data and what happens when agents work with data, and also development workflows. a really exciting talk today. I've known CL for a little over two years, and I think the work he's doing is absolutely fantastic, and I'm excited to bring him on the show here, talk about some of the cool stuff he's building.
CL, welcome to The Task at Hand
[00:00:30] CL Kao: Thank you, Jakob, and it's great to be here
[00:00:32] CL's Background: SVK to Data
[00:00:32] Jakob Heuser: Awesome. Yeah, let's, let's kinda jump in because I think there was a world before, the work you're doing at Recce, I kinda wanna start there and give listeners a bit of background. I've known you for two years, but everyone here might, has known you for about 30 seconds now. let's talk about what happened before Recce, 'cause there was g0v, there was SVK. Walk us through how you got to, "Oh, data's the problem, and this is what we're gonna chase."
[00:00:58] CL Kao: Oh my gosh. so trip down to memory lane. so I started off as a competitive programmer, back in Taiwan, and then I was on the national team, for the international Olympics for Informatics. So basically, you are locked in the room to solve, algorithm problems, when you're a high school student.
and then, so from there, it's actually, traditionally we got taught that, program is, data structure and algorithm, right? So basically, now you have data, you have program processing the data, and then throughout the last couple of decades, it becomes like, program are actually data from beginning 'cause they are instructions, right?
But then, comes machine learning, these classification or anything are actually, constructed by data. And then LLMs basically take us to the extreme, right? But let's start on, the very OG traditional programming. when I was doing a lot of open source work, like Apache or, Subversion at that time, I was frustrated with, like the lack of, better, version control system to manage, branches.
'Cause, y- I often, have parallel tracks of development doing same thing, d- doing, doing, multiple things on the same project. And then at that time, the best version control system is CVS, Perforce or, Subversion was just there. And then, and then there's no good way to do, real branch, or smart merging.
And then if you, live through those years, you will know that, if you have to vendor branch another software project into yours, this is big pain in the ass and then, so I started to think about what would be... Like, if there are 10 times more people working on software, what does collaboration look like, right?
So it's created one of the earliest, decentralized version control system called SVK. It's a decentralized layer on top of subversion, that allows you have a local clone and then track the branching and then allow you to do three-way merge properly, across all your branch or remote or local or with your friends.
so that's that's about two years before Git. and then that's My journey taught me like the first-hand experience, like if you change the way people collaborate, you basically change the whole industry, right? So I was... I've been always like fascinated about that. and then from there, I've been working on a lot of, ways to improve, the collaboration between, like technology worker, like, ML ops or, like Recce is our, later, take on that.
Like, how do we improve like data professionals to collaborate together now we are going to, more code first approach rather than a BI or, clicky, ETL. so it adds on top of the traditional, co-collaboration, but adds very different requirement and different stakeholders
[00:03:34] What Is Data Review?
[00:03:34] Jakob Heuser: Cool. Yeah, that's, that's, that's actually a really nice segue 'cause I think one of the challenges that we're looking at right now with something like Data Recce is this idea that the data needs to be, be reviewed just like the code. And so in the same way SVK built on top of Subversion, you're making a claim that Recce sits on top of the code review, this idea of reviewing the data. And I wanna start there by asking, what does it actually mean to do a data review, and what does it catch that's different than doing, say, like a regular code review?
[00:04:02] CL Kao: Okay. there are, two ways to look at that. So first one is that if you only look at the data, and then regardless of how this data is produced, right? You naturally want to, first version control the data and then see if there are any drifts or anything that's corrected and then as intended, right?
And then now this is a, a very, traditional, probably, work done by bureaucrats, is like making sure the data published are correct, right? Otherwise they're liable. but fast-forward to, these days, and then these are data constructed through a very complex data pipeline, probably aggregating different source, right?
So a lot of the change was introduced in the pipeline or the logic that produced the downstream data, right? So naturally you become... this becomes more software that you review the logic of the, the transformation or aggregation. Are they correct? Are they as intended? but l- that's not enough. You actually have to look at the resulting data to see if it's, like representing what you want it to produce.
and then again, if it's drifted from the previous version or, now that you add a, a, like a modeling l- table there and then, rederiving the different things that's computed from there, now it's different, right? And then in software, we call this a regression, right? something's broken and then...
But it's very hard to do that in data because it's unlike software where we have a lot of practice like, doing unit tests or integration tests or, however, deep you wanna take in the QA process. but for data, it's, it's a lot of eyeballing. There's... and there's lack of like definition of what correctness actually means because a lot of the time the requirement comes in really ambiguous and then, you have to clarify.
There's a lot of assumption, and then you produce something, and then the original business owner say, "Oh no, that's not what I meant." so the problem w- to address the data review is really that the correctness is ephemeral or subjective. How do we make it, as, as, like deterministic as possible so that you actually have a definition of what correctness mean?
[00:06:06] Jakob Heuser: Yeah, I think that's super important because like you said, data's inherently messy. when we have code, code has one version on the inside and then one version on the output side. And we can look at stuff like a migration, and a migration says, okay, it's gonna transform these columns this way. But code review doesn't account for your several million records where some of them may be completely non-standard inputs because who knows where they came from, who knows how they were put in, who knows what's actually in there.
And the only way to validate is what you're saying is you have to actually look at the data itself before and after, almost like a diff of the data itself
[00:06:47] CL Kao: Of the data itself That's correct. And then, so but the, the problem then becomes, we have all the technology that, you can basically materialize all the versions of the data based on, your different version of the logic, right? And then compare them. But then it's gonna be very noisy. You have a bunch of data, and then the source data is also moving or revised or overwritten.
And so what are you actually looking at, is a messy kind of should you even compare this? and then what is... How... It doesn't solve, if you can derive correctness from that
[00:07:20] What Reviewers Actually See
[00:07:20] Jakob Heuser: Yeah, let, let's actually dig into that. So like what does a reviewer actually see? I can't imagine they're seeing a diff, but make it SQL dumps. They've gotta be seeing something a little more user-friendly with that
[00:07:31] CL Kao: That's right. and then a typical, a use case would be that you're changing part of the data pipeline, right? And then, what we built is something that derives the kind of column level lineage for the impact area. And then for those area are, some of them are more important because they feed into further downstream, right?
And then, so you usually would start with like a, a profile comparisons, hey, is the profiling of the, the value like similar or like different? And then if you wanna really check that, you do a, a, row by row value comparison and all that, right? So there are various lens of looking at that. and then it really depends on like the vol- volume of the data, and then what does it really mean, and then is it like close to the end result?
And then, like I wanna get into this, in a bit, nowadays, if we're having the agents making code changes, almost like as messy as data . So how do you look at like what are agents producing and then how that program feeds into the next program, right? And then how the output of the agent feeds in the next...
So it's very similar problem that you have a tangled, a web of this data flow through that, and then this data can be a deterministic, SQL query output or can be something that agent produce, right? and then to me it's a very similar problem. You are managing a messy transformation of things, and then you don't know where to look.
And then because if you look at everything it's too much, and then if you don't look at anything, then it's very risky
[00:08:56] When AI Touches Data
[00:08:56] Jakob Heuser: Yeah, that, that's actually kind of fascinating 'cause you just said, like, your, your headline over on LinkedIn says, "Exploring what breaks when AIs touch your data." And one of the things that I think is really funny about that is I just had one of my agents go in and write a SQL migration, and I think I've just lined myself up to be a Recce customer because I'm now saying, "Okay, I can't go run this on prod." So now how do I actually test whether this data migration is doing what I wanted?
[00:09:32] CL Kao: Yeah. but, for migration, if, you're, like, just, replicating the logic and then without, changing the output, that's actually the easy part because you just have the agent ensure the result is the same, right?
But for ongoing maintenance of your data stack, that a lot of time you have intended change where you actually corrected a error before, right? You actually wanted to aggregate things differently to, represent your new needs, right? So these are different from, you are making a refactoring or you're just migrating A to B, then you're expecting the result to be exactly the same, right?
So th- I think that's one of the easiest one to solve, but we solved that pretty well. And then, But the harder one is really, like, when you actually intentionally change-- wanna change something, how do you verify the intended result? Because you don't, you no longer have a baseline that is y- this only needs to be equivalent to the previous version.
[00:10:25] Jakob Heuser: Yeah. The data, the data pipeline part I imagine gets a lot more complex because you have a constant flow of data. It's not like you get to take a snapshot of large production and do the migration.
[00:10:35] CL Kao: Yeah, like your product analytics. You literally change it in flight. Yeah. luckily, compared to, 10 years ago, the, it is actually a lot easier to do snapshot and then, like a cheap, a copy-on-write clone for things. Many of the vendors support that, right?
we are, like, at a much better place that, if you are serious about data, you can actually, do the correctness, validation a lot better
[00:11:01] Jakob Heuser: So what, what's--
[00:11:02] What Breaks With Agents
[00:11:02] Jakob Heuser: Obviously, you changed your headline to "Exploring What Breaks When AI Agents Touch Your Data." Let's, let's ask the question: What breaks? What, what actually broke where you were like, "Okay, this is actually the... I'm gonna change my headline"?
[00:11:14] CL Kao: This is what I'm gonna say. So what breaks when agents, changes the data or actually change anything, right? obviously the thing, the first thing is really they are very confident about the changes they make, right? So you gotta be very careful about, validating that or challenging the approach or the result.
but as model get better, they oftentimes have very impressive result first time, right? And then the ongoing maintenance or, like keeping something that's, useful and then, like suitable for a, a foundation for other use cases, is a lot harder. so I think you... everyone sees a lot of great demo for like one-shot, coding agent generated thing, even for data analytics, right?
But I think the harder part is that making those things maintainable and then in a, if you're actually even productionizing it, like internally, then you gotta be more careful
[00:12:11] Recce Customer Stories
[00:12:11] Jakob Heuser: So what, what's your aha moment then? Like, obviously you've taken Recce to dozens of customers. Like, you probably have a very memorable story where a customer looks at this and went, "Oh my God, this is everything that I'm trying to deal with."
[00:12:25] CL Kao: Yeah, we... Yeah. Sure. so th- these were like, even before we have a lot of AI integration into, Recce, and then there was a team that's already using the open source version of that, right?
And then they integrated, into their PR. and then of course, the various different level of, experience on their team. So the more, like more... It's experience, All right. So obviously, the team is composed of different seniority, right? and then usually the more senior people on the data team are more exposed to, software engineering practices.
so they care about like CI, they care about all those things that don't break before merging into production because, this data feeds into their financial system or other thing, then they can afford this to be broken. it is not just for just look at a dashboard, right? So if it's any- anything that's serious, you really have to treat it as part of the product and then be careful about that, right?
so they were already using, Recce for, the open source version of Recce for a while. And then, so when we were like in touch with, "Hey, are you interested in the cloud version where, we just make sure every PR, gets, through that and we also have an AI, summary that's very targeted to, your use case and so on."
And then it's "Hey, CEO, you don't have to sell me this. We already saw it because the one time we didn't use, the open source thing and then things blew up like horribly. Like we had to clean up all this like mess after this thing goes in production, explain to multiple stakeholder why this, was horribly wrong for the past week and et- et cetera."
so that's when I felt, oh, this is actually, valuable. and then people are getting, like useful value from that
[00:14:07] Blast Radius Visualization
[00:14:07] Jakob Heuser: So people can actually go grab Recce today. They can run the open source version of it on their code and just start seeing that value then
[00:14:15] CL Kao: That's right, yeah. So usually it's, like you have a, say, dbt pipeline where, you manage, data transformation, in a code-first way, and then you probably already have a s- kind of semi-CI process, and then this just add on to that for, a lot more confidence then for the team.
[00:14:32] Jakob Heuser: Nice. so if somebody wanted to get started, what's the first big breakthrough they have? They drop Recce onto their product. They start, they hook it up into, like, a GitHub action or otherwise in their CI pipeline.
What's the first thing they see when it starts to collect
[00:14:45] CL Kao: Oh, the first thing they see is their, their CI shows a impact radius for the blast radius for, you're making change to this part of the DAG, and then, these are all the downstream impact and then, in a very granular way that you have a column level, lineage, and then this column level traverse in different ways, right?
If you're just passing through the data or y- you're using that as an aggregation key, these are all have different implication, right? So you really see a very concise, a visualization for "Hey, I'm making this change, now I know this actually has a, a, a huge impact a- and I should pay more attention to that," right?
So it's inherent, a risk classification, right? If you know that this touch a very important, data model, then, you should pay more attention
[00:15:30] Jakob Heuser: No, you used the term blast radius, which I think is a really good term. The, the term blast radius where it's like, well, I just changed one column, or I just changed one... I changed one, one table view, and I didn't realize what other systems depended on that data looking a certain way.
[00:15:49] CL Kao: Yeah. Yeah
[00:15:50] Jakob Heuser: for a lot of people, that's gonna be their first aha moment.
They're gonna be like, "Oh my God, I didn't realize this change was as big as I thought it was
[00:16:00] CL Kao: That is exactly right. And then, we have that, in, traditional software, you have code graph, and then usually these are all pretty clear before you even, commit or, make a PR, right? But, in the case of data or even in the case of, if you're, having some agentic system and then some prompt are version controlled, it's basically just like data.
You don't know who's consuming that, right? But you could, but you have to be very, conscious about, this is being used by certain other system and then what does it mean that if we change something here?
[00:16:32] Custom Rules & Risk
[00:16:32] Jakob Heuser: So what's the extension of Recce look like then? Obviously, Data Recce's already helping people with their data problems, but there's work still to be done because I think the other half of it is not only the blast radius of here's all the consumers, but there's a code piece as well of now I've gone and I've looked at all your consumers for you, and I can now quantify how much work this change actually still has left to make.
[00:16:39] CL Kao: Yeah, you're right. And then, for the actually impacted column or the value, right? So you are able to explore that to see whether this is intentional or this is a accidental change that you didn't intend it to, right? And then you can create custom rules that, hey, every time I touch this thing, this is so important, just show a very detailed, row-by-row comparison, or, this is not so important, but I wanna make sure the aggregated result is the same, right?
So you can Depending on how, important one data model is being used, and then you can customize the rule that you wanted to see. It's basically if we think about, the traditional PR being, about code, right? Let's think about the extension to that being like, what does correctness mean, right?
In the data sense, like what evidence do you need to, to be comfortable that this is correct, right? so again, you... If you generalize the problem, this is the same as you have agent do the work, and then, and then you s- you clarify the goal and then... So imagine now this is a very narrow case for the agent's making change to the data stack, and then, and then it's not like some other, some people on your team.
What do you need to verify, the result from the agent? so it's essentially the same question, right? What would you ask to this PR that, to prove that the resulting data is correct?
[00:18:11] Spacedock: Agent Workflows
[00:18:11] Jakob Heuser: I, I like your fixation on correctness because I... is kind of a nice segue into something that I think is one of the coolest prod- products that's come out of the Recce journey then. you're also the creator of Spacedock. for those that don't know, Spacedock's tagline, "The bottleneck is judgment, not generation." Spacedock basically is a product, a tool built out of the team at Recce that says if code generation is free, then correctness is the cost. I'm gonna let you talk about that a little bit because I think we have to start with... I'd love for you to say a little bit about why that is.
[00:18:50] CL Kao: Yeah, absolutely. And then, I think, just like many, people, I think we got this, the, what is it called? the, the twenty twenty-five, Christmas break, like agent moment of wow, this got really good, right? and then, and then this is no longer, vibe coding. this is actually a... or s- we'll call it agent engineering nowadays, right? So you have a way to work with very capable, interns that are probably ADHD, have no taste, and then, but they are very smart, right?
So if you give them good guidance, they can be very effective, right? and then so I was, using those agents to, do some side project or improving and then all that. And then I got this like, this kind of, thrill. This like, wow, I'm like on top of five different things, and then I can just give direction and then things will, materialize at the end.
I can play with, I get feedback, or I construct a loop that, the agent observe the result and then correct that, in, in, in some way, right? So it's like a early version of Ralph Loop. And then, and then so that was the time, like you have a two day of, this like ecstasy. It's oh, wow, this is awesome.
And then you're super exhausted 'cause you know what? You're context switching between ten different things at different levels, right? You have the project planning thing or like implementation thing or data validation thing, right? And it's got very different type of exhaustions. oh my gosh, I can't do this.
[00:20:22] Agent Exhaustion & Context
[00:20:22] Jakob Heuser: I, I know, I know it's not on the question list, but that exhaustion is something I think we all feel when we're working with agents. And I think
[00:20:29] CL Kao: Yeah
[00:20:29] Jakob Heuser: I think you just put a finger on it, that it's the context switch. It's sure, I'm getting a, I'm getting seven or eight concurrent things done, but every time I switch tabs, I have to basically be like, "Wait, what was this
[00:20:43] CL Kao: What am I even doing?
[00:20:44] Jakob Heuser: am I, what am I doing? Where am, where am I?" And it soon becomes a when am I? And by the time 9:00 PM rolls around, you're just like, "I don't wanna think anymore."
[00:20:53] CL Kao: Exactly. Exactly. So I think everyone, a lot of people I guess I know, went through that progression, especially very experienced, developer from the past. and then they all get is like, "Oh my gosh, I can be like, in my 20s again, like coding all night," right? But it's a very different type of exhaustion.
and then at that time I was using, Jesse Vincent's, superpowers. and then it's a great, scaffold for, the agents to do the plan, do the, development, validation. and then I then very quickly realized I'm doing the same thing. I am instructing the user... the agent to, do the plan.
And then, I review some high level thing. I have another agent reviewing that, and then implementation, and then validation and so on, right? So there's gonna be a better way that, I think at that time the agent are, still not super smart, so those scaffolding help a lot. But then I started to realize, what actually helps is actually, for, for the human to understand, like, where things are, I have this idea for this, project to go a certain way, and then where is it? is it in planning? is it being done? Is it being validated? Have I tested it, right? So it's actually helping me. And then I did a benchmark on, agent working on, some data problems, right? Solving data problem.
And then, and then using superpower, and then it actually help the weaker model. It doesn't quite help, the stronger- strongest model. and then there, there's several, other research on that as well. I believe that, for a real smart model, they don't really need that scaffolding for, you gotta write the plan first, you gotta review that, blah, blah, blah.
They can figure it out, right? But it actually still helps human, like, where things are. Like, can I test it now? Is this still a prototype or is this kind of production ready?
[00:22:33] Jakob Heuser: Yeah, the, the,
[00:22:34] CL Kao: so anyway, yeah
[00:22:35] Jakob Heuser: that, hey, there's seven steps in this plan, and where you're jumping in and where the, where the human gets b- back involved, we've done three of the seven. So you know where you're at in the sequence. And sure, the agent holds all that in its memory, but if there's no artifact, you're jumping in and now you're trying to figure out, looking at the log, where did it pick up?
Where has it left off? Where am I actually needed?
[00:22:59] CL Kao: Yeah, you're right. And then so it is like that, exhaustion plus this, frustration is like, why am I like, being, becoming the human connector? Am I, am I being looped by the agent or am I operating the agent, It's like I kept being asked like, "Hey, can I... Can you approve this? Can you approve that?"
and then by the time at, in the evening, you don't really have the energy to actually read through that, but you wanna keep going, right? Everyone's done that, just... and then so I was like, "Huh, there's gotta be a calmer way to work with, these like ADHD, interns," right?
so that's why, Spacedock was born. It is, a, a way to declare, your workflow in a lightweight way that is basically a DSL in a YAML file. sorry, in a Markdown files from Matter. and then so you describe how you want to instruct the agent to go through a certain stages, and then what's your relationship with the agent, at each stage, right?
Do you want to review the plan or, do you trust the agent enough that the plan can be auto-approved? so this trust obviously builds over time, right? So when you're comfortable with your workflow that, and then maybe you can start running a sprint. "Hey, I have five different tasks and they are all like connected."
And then, but figure out a best dispatch way, to the order of dispatch to finish that, right? so that was my first, aha moment once like the very first prototype of Spacedock worked, right? I started to, emulate what we always do in, software development team in sprint, right?
let's plan the sprint, have everything like planned now and then review, right? And then now go off like dispatch this like as efficient as possible, but still do the verification at the end, right? and then I went to bed, and then the next morning was like, wow, there's a 10 PR ready to merge.
It was like, oh, wow, this is awesome. and then so basically I started from there. It's like there, there are interesting way that, we should think about our relationship with the agent. Like when... and then it's very workflow specific. It is... it, a lot of the time, like a, a prototype project, you don't really need all that, right?
But as proto like graduate into something you wanna maintain, if it's own, your own project or if you're contributing to other people's project, they are all different workflow, right? so all those things, can you uncover that? Can you decide your relationship with the agent, right? I wanna review all the, the kind of important architectural plan or I wanna review it with, another expert agent's, comment so I can make informed decision, right?
So I think it's about that dial that you want for, how you wanted your relationship with the agent to be that is the core, thesis for Spacedock.
[00:25:35] Managing Your Attention
[00:25:35] Jakob Heuser: It's, it's very interesting 'cause it's this idea of managing your attention. as an engineer, if, if you've got-- If you spend 20 minutes on a context switch and you're gonna spend 20 minutes on an individual agent's output, 40 minutes total, and to get the value out of it, what should you look at?
And I think that's a really hard question. When an agent has done the plan, has done the architecture, has done the diagrams, has done the code, has put everything through a bunch of gauntlets of tests, including, I'm hoping, rec- Recce, and then becomes: Where do you as the engineer, as the architect, as the operator, where do you spend your time looking at the collective output? that's-- Previously, it used to be you spend your time on everything, then you lazily enter your way through it. But technically, when you were saying accept, accept, accept, accept, accept, you were building trust. You were saying, "Yeah, I do trust it at this point," weren't codifying that trust anywhere.
You weren't saying,
[00:26:34] CL Kao: Exactly
[00:26:35] Jakob Heuser: okay, you, you're good at writing plans. We're, we're good at the plan part. loop me in later when you get to the 'I found something that would b- be a backwards incompatibility problem or would force a major version bump.'" Like, okay, loop me in then. And those, that terminology, that jargon or that system didn't exist before a tool like Spacedock it was ask all the time or never ask at all.
There was never the judgment piece. When should an agent ask for help, and when should a human get involved? And your workflows in Spacedock YAML files really describe lot of that, like, really hard piece right there, which is: When is a human actually needed?
[00:27:16] CL Kao: That is right. And then I think a lot of us like start off with, like very cautious with the agent. You probably approve all the permissions and then gradually it's oh, it's dangerous to skip, skip permission or auto mode, right? But these are like mechanical, or I would say like syntactic decision, right?
These are like, can I do this? Can I do that, right? But a higher level semantic decision is more like, should we go this approach or that approach, right? It's almost like when you're running a team, then, then you want your most senior engineer to talk to each other and then give you a plan that you can say yes to, right?
As opposed to y- because they are in, on the problem. They're smarter than you, and then you shouldn't go diving in to interrogate every detail, right? You should trust them. And then, but you got to build that trust first, like by having a way that, like what does Jakob expect?
What does CEO expect? What would CEO ask, right? And then have those like resolved, before they get even presented to you
[00:28:10] Jakob Heuser: Yeah, because your goal as a, usually in a manager or a tech lead role is not what are we doing, but did we actually consider everything? Like, I,
[00:28:19] CL Kao: Yeah
[00:28:20] Jakob Heuser: I at, at some point as a manager when I was managing teams at Pinterest, it wasn't about jumping in on the technical decisions, it was does their plan look complete?
Does their plan look like it's thought about where it's gonna run into rough edges and where it's gonna fail? It wasn't about whether or not it was technically correct. I,
[00:28:37] CL Kao: Yeah. Yeah. Yeah.
[00:28:38] Jakob Heuser: in, like, forever. Like, I'm, I'm not gonna know if this is the right way to do it.
[00:28:43] CL Kao: Yeah.
[00:28:43] Jakob Heuser: not the question I'm supposed to be answering
[00:28:45] CL Kao: Yeah. Yeah. You're exactly right. and then that's how, I think, everyone is gonna be managing team of agent and they gotta learn this, management skill. It's not about diving into all the technical detail. And then, but you still have to, figure out the nice or the, the correct choke point that, the best...
the question you asked ear- earlier. I would add to that, though, is, can this be simpler?
[00:29:06] Spacedock at Recce
[00:29:06] Jakob Heuser: Yeah, it, it, and I think it's a really hard problem. Like, you, you obviously are using Spacedock at Recce. what has the team been able to do with a tool like Spacedock that wasn't possible before? Like, how, how has this
[00:29:20] CL Kao: I,
[00:29:20] Jakob Heuser: build stuff?
[00:29:22] CL Kao: so before using Spacedock, the team are actually using a combination of different things, like, we- our own, review loop for code change or our own kind of data, agent for using that. So Spacedock standardized all that and then now you can have a concrete goal, like, "Hey, I wanna burn down, this 10 minor UX issue."
Go do that, right? and then it will follow a certain standard, follow a certain thing, and then all the predefined, criteria that they should reach out, to human for help. so before that, we actually have so people... Some, some team member use the Linear agent or other approach or the home brew thing, right?
and then it also adds a little bit of, cognitive overload for other people to take over the wor- the work. so it's standardized a little bit, for the team and then a- across different workflow, because you use the same terminology, you use the same is this stage, gated, as in needs human approval or not, right?
and then the flexibility is really, once you have the workflow, you, trust that this has got, like, all the, the important thing, predefined, then you can say to the agent, you have the con to execute toward the goal for all those different tasks, and then just make it happen.
And then, like you, you are authorized to merge and, run the final report." So it basically turned that, the building block in, a, a given point of time that a decision needs to be made, right? And then you can predefine how the decision should be made and then authorize the agent to do it.
[00:30:50] Jakob Heuser: And then,
[00:30:50] Lure Fixtures & Over-Engineering
[00:30:50] Jakob Heuser: recently I, I saw in 0.26 you rolled out lures. For those that don't know, the lure system is basically fixtures. They're honeypots designed to trick agents, I mean that in a good way because one of the things that we've noticed in this latest generation is Codex, like if you're using Sol or you're using Opus 5, Fable, they love to over-engineer.
They love
[00:31:12] CL Kao: Oh my gosh.
[00:31:13] Jakob Heuser: they prema- they prematurely optimize. and that ends up being a maintenance headache. And so you actually built these lures, these, like, tricks basically, that trip gates inside of Spacedock. If it runs off and it starts working on a bigger problem than what it was, it drags a human back into the loop. Talk to me about, talk to me about how that works, because I think that's not even a speed bump. I think you've described a new primitive of, like, sensors in these workflows of, hey, you tripped the sensor. A human is needed because you're, you're going in a wrong direction
[00:31:49] CL Kao: Yeah. Yeah, you're right. and then this is exactly the frustration, from, I think Sol and Opus 5 specifically, are particularly, prone to, solving very pedantic problem. so the, so the lure, the honeypot was, initially designed to see, are we able to mitigate that, right?
'Cause, you can give... I know Taskless is, giving rules to the agent, like, and then if you say, "Hey, never build a test case for the test suite," right? you test the thing rather than, build this, the test to test the test and the test, right? And then, you're laughing, so you've run into that.
And then all these, pathological tests, right? So we can give, instruction to prevent that, but how do we know they are effective, right? So the lure was actually the specific scenario where, that it described the scenario, the, the user wants to do a certain thing and then the plan doc has already documented and then the test is required, right?
Is the agent, gonna go off, build a whole, So one of my worst, example was, I was building, a terminal markdown, reader, as part of Spacedock. It's called Subspace. And then one of the test suite, that the agent decided that instead of using tmux, it's gonna build a whole terminal emulator, to actually test the behavior for rendering and so on.
I was like, no. Stop." and then I think that, that was actually, what inspired the, the honeypot, project is like, "Hey, given this case, we are building a terminal, markdown file, and then the instruction says we need to test it. What would you do? Would you build a terminal emulator or you decide to use tmux, right?"
So basically, it's basically a test suite to, to eval whether, The agent get lured into overbuild or to care about, what if this file is a symlink? And then what if, other people can write to this file? No, I don't care. This is prototype, right? so one of the thing, of course, is you have to specify the context for this current project being a prototype.
You don't have to care about all those symlink bullshit, right? and then, But anyway, so I was, like, deploying a quite a few, changes to the workflow and what it looks like for, for our system so that this can be prevented. So those honeypot were actually, to verify whether this, mitigation work or not.
and then later we figure out, this is actually a important thing because just like changing data pipeline has impact to the resulting data that could be very bad. Changing, agent instruction and then, having the agent behave differently might not be what you mean as well, right?
we create a tiny tool called Behavior Diff that, whenever you modify or you intend to modify a skill or a CLAUDE.md, AGENTS.md, it's gonna synthesize what are the potential, scenario that, that kind of, inspire you to do that change. What was going on before, you wanted to make a change?
And then if we can synthesize a couple similar scenario, does the change actually fix that, right? so this is a hard problem because it's all non-deterministic, but we gotta start somewhere. and then oftentimes you don't want to curate a full eval suite for a tiny, change in your skill file, right?
So this is one of the, the, the other, tiny bo- bit of, space that we built in to visualize the behavior change o- once you change the definition of a workflow.
[00:35:02] Jakob Heuser: it is a bit, like a lightweight way to construct an eval set.
[00:35:06] CL Kao: " Hey, can I ensure my agent's behavior is this way?" so that when you make changes to, whatever is a skill or your workflow that, might change the agent's behavior, you have a better confidence about that where, before you build out a whole eval suite. so that actually becomes, a part of, Spacedock as, as a behavior diff, like, whenever the agent propose a change to workflow or, that you wanna manually change the rules, and then it will synthesize, relevant scenarios that, that might be, relev- that might be, like, why you wanted to change the behavior and then tell you the before and after in a visual way that, in this point, the agent would decide to do this, but your intention is that, right?
So does it actually ref- reflect your, preference for the agent's behavior?
[00:35:48] Jakob Heuser: That's awesome because, I mean, when code is free, ultimately, like at Data Recce, you're gonna be doing a ton of prototyping. you're gonna be prototyping all kinds of ideas, and in those situations you want the agent to move quickly and not try and boil the ocean. And then some percentage of those ideas are worth doing that second pass investment in, in those scenarios, now you want a little bit more overbuilding. You probably still don't want it to rebuild tmux, but you probably want like at least a little bit of like, "Okay, yeah, do consider a lot more edge cases because now we're talking about a production workflow." And it sounds like Spacedock can do both now, and it understands if you set it up for doing prototyping, you can keep it towards very focused, get something in front of a couple customers to be like, "Is this what you wanted?" rolling out the actual production solution
[00:36:38] CL Kao: Yeah, exactly. And then if you're familiar with all those kind of, s- spec-driven development, there are plenty of frameworks to do that, right? And then, but the problem is usually it tries to instill a, a pretty rigid, workflow that, either is planned and then implement review or some other, flavor of that, right?
and then my bet is really, we gotta be more, a lot more flexible but repeatable for different type of workflow. for inspiration, right? You want a, a very quick prototype, and then you care about certain thing, before a production, or making, a proper change or for a data workflow, that you're trying to answer a question and then trying to decide whether your core modeling is changing or this is a one-off thing.
So these are all different workflows, but, they need to be managed in a way that, the human and the agent col- can collaborate meaningfully.
[00:37:27] Productionizing Agent Ideas
[00:37:27] Jakob Heuser: Yeah, I think, I think that's gonna be sort of this next step is now that, now that we're at a point where can generate a ton of code, we can generate a ton of product ideas, the idea of how do you productionize one of these ideas is this entire workflow will be a problem
[00:37:45] CL Kao: Oh yeah
[00:37:46] Jakob Heuser: solve. Like, is it a re-implement problem?
Is it a migrate and take what we have and basically change the way it behaves? It's gonna be different in every company 'cause every company's gonna prototype stuff in fundamentally different ways
[00:37:58] CL Kao: You're exactly right. And then, that's, like why we built Spacedock
[00:38:03] Jakob Heuser: so the team is on Spacedock over at Recce. what did they have to give up? Did they have to say goodbye to their Linear agents, their custom workflows? Like, what did they have to give up to sort of move into the Spacedock world?
[00:38:16] CL Kao: so I think it's less spec- specific, but more how we see the future of, agent-augmented, engineering, look like. And then, we touched on this earlier that I think everyone's gonna be running a team of agents. y- if you're IC, you are still, like, having a managing job now, congrats.
and then, and then... But, th- the good and bad because y- you don't need to deal with humans, but you deal with a bunch of agents that are, very s- again, very smart and then, w- ADHD without any taste. And then, so your ability to guide them through a better result, through, what type of workflow that you're comfortable with given this, your context, is super important.
And then I think we, we can l- actually learn a lot from how we run software team, to borrow that into running agent teams. I'll give you an example. we talked about Sprint earlier, right? it is actually, pretty effective, to borrow the Sprint ideas that you have relevant feature you group together and then have all the plan and review cross-checked, and then you set off to development.
One thing that's different from traditional Sprint is, we were, using this, the Sprint approach in development team because we have fixed capacity. and then so sometimes it's a release schedule. You have a certain theme, and then you want also other bug fixes to be here in the Sprint.
But I think for agent, it's unlimited. you can afford a token. it no longer makes sense to, to, to cap, your, like your develop- your delivery, capacity, right? but it still makes sense to group similar feature, in, in terms of what you want to deliver.
but you can leave out kind of all the bug fixes. It can go off as a separate agent. so I think there's a lot of things that we can learn and then, but should relearn from how we run software team and to be, more effective working with agents.
[00:39:56] Jakob Heuser: No, that's, that's actually really cool because I, I don't wanna use the term manager 'cause we're gonna scare off all the builders and makers that specifically chose to not be managers. but I think you're right in that we are all becoming tech leads, whether or not we want to. you're not necessarily always the person now writing every bit of code, you are the person, responsible for holding the system in your head about how all these pieces work together. And tools like Spacedock help codify how exactly am I going to... They're, they're codifying how do I jump in and look at things? And then
[00:40:28] CL Kao: Yeah, that's right
[00:40:30] AI Surprises & Sandboxing
[00:40:30] Jakob Heuser: And tools exist for when you see something, you're like, "Yeah, we don't wanna ever do that again." how do you build that enforcement layer so that the next time an agent comes in, it doesn't bump into those and bother you with having done the wrong thing again? some closing questions on the way out here. One of the ones I love to ask everybody that is in the AI space, is just two questions. What is a capability that AI has h- has now that surprised you? What's something you didn't think it would be able to do?
[00:40:58] CL Kao: there, it's always been a step function, right? So I can, I guess Star was my first, wow moment. it is, when I was doing a pretty complex, rebase, for a branch, and then the agent say, "Hey, I tried a couple times, but I would actually port the change, independently rather than try to, force rebase because it's too complex."
And then, so that was I think almost a year ago now, and I was like, wow, okay, this is 'cause I worked on version control system. I know you can get stuck into a merge conflict and then how do we rebase this thing properly, keep history. And the agent was able to break from that specific instruction to say, "You know what?
rebase is not the most effective way to do this. we should port this change over." So this is the first one. The other one is, there was an- another prototype I was doing and then, and then I powered through that to get a first, working demo for a, it's like an agent log aggregator.
And then I asked the agent, "Hey, what have we learned, throughout this journey, like, building this thing? And then how... what would you do differently if you're starting from scratch?" And it lists out a lot of like learnings and then I was like, "Huh, you know what? Actually, go re- restart, doing this, in this way you approach."
And then that was the first, long running session. I had about seven hours of like reimplementing all th- all the things, because it, it had learned that all those, edge cases and, So it recreated a plan for a full implementation of that system. so that was my second one. and then I think this is But coming back to like how this kind of weird world we're in now, it's really, essentially everyone's having a botnet, connected. Like you're allowing agents to do a lot of things, on your behalf, in your, system. so if you don't control that, it's really dangerous and it is really, like alarming, with all this, OpenAI and Hugging Face, incident, right?
And then this can happen because we are allowing the agent to control a lot of devices, right? So I think there's a lot of how do you wanna control workflow? How do you wanna control sandbox? And then what is, what needs approval? I think these all need to be, like reconsider and then we need new, construct for all that.
[00:43:05] Jakob Heuser: Yeah, sand- sandboxing is something, yeah, on local machines, I don't know if you run dev containers or anything like that. I've run a ba- bare metal a lot, and I have a huge approval list of things that are both allowed and denied. I know I'm running, I'm not running dev containers, I'm not running in a Docker environment,
[00:43:20] CL Kao: Yeah.
[00:43:21] Jakob Heuser: it, it could at any moment in time decide to run a shell command, and going and getting a new laptop kind of thing
[00:43:28] CL Kao: yeah. There, there are a lot of I use a very lightweight sandbox called, Agent Safehouse. So it's a, sandbox exec-based one on Mac.
[00:43:36] Jakob Heuser: Okay
[00:43:36] CL Kao: so you have to be explicit, like what is the agent allowed to do? a-accessing certain file or using system services or something like that, right?
and then on VM, I think you are like, m- less restrictive. It's yeah, just go YOLO.
[00:43:52] Jakob Heuser: Yeah. No, A- Agent Safehouse is really good. We'll put that in the show notes as well. finally, where, where can they find you? Obviously, like, you are, you're on socials. there's Re- there's Recce. We'll put a link to Recce in the show notes as well, along with a link to Spacedock. But where can they find you if they want...
They're like, "The, this CL guy, he knows his data shit. He is very, very knowledgeable when it comes to agentic workflows. I wanna follow his stuff. I wanna know what he's up to. I wanna, I wanna stay informed when he comes up with new stuff on Spacedock." Where can they find you?
[00:44:19] CL Kao: yeah, so you can follow me on LinkedIn and, there's a, a contact form on spacedock.md, and that, you can just put your email. We'll send updates there about, this, new agentic, frontier that we're exploring.
[00:44:32] Jakob Heuser: Okay, so primarily LinkedIn. I'll g- I'll get the links in there. I'll get everyone connected to you. everybody that's been listening today, this is CL Kao, creator of Spacedock, founder of Data Recce, AKA Recce. everyone give him a follow. Everyone check out what he's doing, because data's gonna be a lot better because he's involved. I'm Jakob Heuser, CTO, co-founder of The Task at Hand. We'll see you next time when we talk to the builders and makers and the stuff they put into the world. Take care, everybody
[00:44:57] CL Kao: Thanks for having me