September 27, 2026
-
Another question I have is how exactly to launch agents. In other words, how to turn a description into a set of running agents, either in a k3s cluster or in microsandbox.
Suppose we already have a description of the agent team.
Now we need to bring the cluster into the state that matches this description.
I see two approaches here.
The first is static.
We run
run. It takes the description and generates a regular set of Kubernetes manifests.We can inspect the result, verify it, put it in a repo, and then let a GitOps controller apply this state to the cluster.
The second option is much cleaner architecturally.
We could build our own Kubernetes controller that reads the description directly and continuously reconciles the actual state of the cluster with the desired state.
But right now, I do not think we need to start with the second option.
The static option is much easier to understand.
Run
run.Get a concrete result.
See which pods and other objects should appear.
Check exactly what the system is going to do.
After that, if it becomes clear that we really need a continuous reconciliation loop, we can move the same model into a controller.
So my next step is not to "write an orchestrator."
The next step is much simpler.
First, I need to describe the target state of the system after
run.What exactly do we want to see in the cluster?
Then we can decide which mechanism is best for bringing the cluster into that state.
First, the desired state.
Then, reconciliation.
-
At first, I thought of an agent description mainly as build instructions. There is a description. We use it to build the agent image. The image includes the environment, tools, and everything else the agent needs to work.
But now it looks like the description is needed for much more than the build.
In fact, we have at least three operations:
build
We build the agent image from the description fully automatically.
onboard
We need to create all the external entities through which the agent will exist in the system.
For example, if the communication layer is Telegram, we need to create a bot with the right name. If it is Zulip or some internal tool, it will have its own equivalent.
This is where issues can arise that cannot be properly handled during the build. For example, the name may already be taken. Or we may need to create a secret and put it in secret storage.
So onboarding looks more like an operator task.
run
The same description is needed here again.
An image alone is not enough to run the agent.
We need to know what the agent is allowed to do, what policies it has, which secrets to connect, what resources it needs, and what state the infrastructure should be in after launch.
So the description ends up serving as instructions for several systems at once.
For the builder.
For onboarding.
For the orchestrator.
This changes how I think about the agent description language.
It is no longer a config for a Docker image.
It is the source of truth from which a working system is created.
-
I am thinking about the agent session as an object and the level at which it should exist.
There are two options.
First, a session is a harness-level object. There is one permanently running agent instance, and parallel sessions are created inside it.
Second, a session is an infrastructure-level object. Each session gets a separate pod or sandbox and effectively becomes a separate agent instance.
The first option is easier to implement, but the more I think about it, the more I prefer the second one.
If sessions live inside the harness, the harness suddenly has to do a lot. It has to create sessions, manage them, restore them, and handle concurrency.
This also creates an unpleasant coupling. If the agent pod goes down, all the sessions running inside it may go down with it.
If each session is a separate infrastructure object, everything becomes much simpler.
A session is a fully isolated environment. It has its own environment, repositories, variables, and possibly secrets.
If one session fails, the others do not notice at all.
Most agents do nothing most of the time. If an agent instance is created only while it is working, we do not need to keep resources allocated for every role all the time.
We may have a thousand or 50 thousand defined roles. That does not mean we need to keep 50 thousand pods running.
Resources appear only when an agent is actually working.
Of course, orchestration then has to move somewhere else. We need a separate layer that can create these agent instances and manage their lifecycle.
I do not yet fully understand what it should look like.
But the principle itself now seems quite clear to me:
an agent session is an infrastructure object, not an internal harness abstraction.
-
Watching how my agents work, I have arrived at a concept I call the Boss session. This session does no hands-on work. It only watches all the other workers and gives them instructions in natural language.
This may look like using subagents, but it is actually the opposite approach.
Unlike subagents, workers are autonomous units. They have their own context, access, and sessions.
Workers listen to the Boss and interpret its commands based on their own context. They do what their role requires and publish a brief result of their work, reporting it to the Boss. This removes all micromanagement from the Boss and lets it focus on the work. It implements the concept I described here: https://andysmith.ai/2026/Sep/25/separating-core-work-from-technical-work/.
The Boss session uses the most expensive and best model available, so every token is very expensive. The workhorse models are chosen for the task, and their sessions are short-lived.
Ideally, this approach will maximize token savings. This is very relevant now. Even though Opus 5.5 is quite efficient, I think the limits will be cut very soon, and it will be very painful, as it was in early September.
September 26, 2026
-
I am reading Navalmanack with my own eyes, without agents. I am now at the part about finding and endlessly refining my Special Knowledge.
I realized that social networks such as X/Bluseky/Threads/Instagram/Telegram are a good tool for finding and confirming that Special Knowledge.
Every social media post reflects my experience, a unique step that sets me apart from everyone else. Together, these steps form my unique path from birth to where I am now.
So every post should be more than just a thought or an idea. It should confirm something I have done in the real world. It should reflect a change, an improvement, something that gives me experience.
So that when I scroll through my feed many years later, I can say, "Wow, did I do that? And that too?"
Through the lens of these small steps, small wins, and small changes, I will find what, as Naval puts it, I am the best in the world at. It must also be something that can be scaled and turned into a product, because if it is only interesting to me, it is not Specific Knowledge.
Once I find it, the best place to put it is in the Headline, so that everyone, especially me, can see it. And, of course, I will refine it regularly.
-
A simple, elegant format for drawing UX components in text. It can be used in reports and documentation, and probably in TUI applications.
-
An open plain-text format for storing health, workout, and sleep data could be a good alternative to Apple Health. It would avoid vendor lock-in and be easy to integrate with agents.
ledger/hledger and their transaction format could serve as a reference, adapted for health data.
It could store both aggregated data and raw metrics.
The format should be simple enough to enter data manually, through agents, or with basic exporters. At the same time, it should be flexible enough to store different types of data.
It should also support machine validation so that reports and aggregations can be generated automatically.
I wonder what ready-made options already exist in this area?
-
I can view the space of possible options as a graph, and research aimed at solving a problem as finding a path through that graph.
At any given time, there is a set of visited vertices. Following either the WFS or DFS algorithm, I can choose the next vertex to visit, which is the next question to explore.
Breadth-first search performs better in terms of speed and number of steps.
But people, including me, still prefer depth-first search because it feels more natural.
This observation is counterintuitive, but it is supported by graph theory. So I need to stop myself whenever I want to go deeper instead of broader.
September 25, 2026
-
Very often, instead of doing useful work, my agents get distracted by technical tasks: creating issues, pull requests, and commits. This fills up the context and reduces the efficiency, and therefore the quality, of their work.
I have tried different approaches. For example, I moved technical tasks into a separate tool that the agent can call. All the complexity stays inside the tool, and the agent only gets the result. But this does not solve the problem because the agent still needs to "think" about calling the tool, which takes attention away from the main flow of work.
It seems that the only option is to move technical tasks one level above the agent. In other words, an external layer, which I think will be oml.sh, runs outside the agent, prepares the environment for it, monitors it through ACP, and handles technical tasks in a separate session.
This lets the agent focus on its core work, while the environment handles the setup and related tasks.
-
There is the good old KISS principle, which I always forget. As a result, I end up with overcomplicated solutions that solve problems that do not exist yet.
How do I deal with this? Especially in the age of agents, when adding bloat is free, but cutting it back is expensive?
Am I the only one with this problem, and need to fix myself, or is it a widespread issue, and we need to fix our processes?
-
I tried to build smith.wiki as the single llmwiki, aka Zettelkasten, for all my research. Within just two days, it turned into a mess that was impossible to navigate.
I think this is because an agent writes far more than a person, or even another agent, can read and review. Garbage multiplies on an unprecedented scale.
Now I am trying a new setup. I publish each independent research thread as a repository in the smith-wiki GitHub organization. Since the CNAME is set in the main repository, all repositories are automatically published at smith.wiki/
. A repository can be public or private. The old setup did not allow private research. The main index only provides a list of research projects and possibly a general index, such as a list of pages in each project. The research projects themselves are autonomous and can use any structure.
September 24, 2026
-
I plan to try running Jev inside EMACS so that it predicts the next command to execute.
The input is the screen state + task, along with the possible next commands. The expected output is the probability of the next command.
Since every action in EMACS is a command, Jev can do anything this way, absolutely anything.
The only problem is that there are an incredible number of commands, and it is unclear how Jev will handle this.
This might be optimized somehow. For example, the commands could be organized into a tree, first selecting a group of commands and then the right command within that group.
September 23, 2026
-
What if, after entering the wrong password, a hacker saw seemingly valid, but outdated, data?
The authentication process currently includes a quick check. You only need to enter the wrong password to learn that it is wrong.
What if every password were correct?
What if, after entering the wrong password, a hacker saw documents, emails, and data generated by AI and unique to each password?
This would make hacking much harder, especially when the hacker does not know what to look for and therefore cannot quickly verify whether the password is correct.
-
Claude Opus 5.5 and GPT 6 Sol are out.
Claude has already appeared in omp for me, even though I have not updated it in a long time, while openai-codex still has no updates.
They say they are cheaper and better now, but I will test them myself tomorrow.
In recent weeks, I have really noticed the lower limits. It turns out that the 50% free allowance over the summer made a big difference. Or maybe I have started working more.
So the lower token prices come at just the right time.
-
To be honest, “sending an email to a blog” sounds very outdated.
Straight out of the last century.
Still, I have a hypothesis that email is an undeservedly forgotten communication channel, and a very convenient and well-designed one.
I’ll try it for blog posts, and then maybe move all communication with agents to email.
-
I’m trying to test my new publishing workflow.
I sent this post by email.
September 19, 2026
-
The idea is that the questions I ask may be interesting to others, not just me.
I developed this process: I post a question in a dedicated Zulip chat. A researcher picks it up, publishes the question, and starts researching it. The researcher then gives a short answer and links to the full research.
I chose Threads as the publishing platform because I thought it allowed free publishing through its API, unlike Twitter.
My agent was blocked after its third automatically published post :)
So I realized that I should not do this on centralized social media platforms.
I’ll try Bluesky.
September 16, 2026
-
https://github.com/coldteadotai/pr-lens
We could use a similar approach to visualize processes in Zeno.
We could also show different levels of nesting, depending on whether you need to dig deep for debugging or just take a quick look at the processes.
This could run as a web service inside Zeno that starts up automatically.
September 15, 2026
-
A small, fast harness. You can use it as an alternative to the big omp inside Zeno.
It supports ACP, so it should just work.
September 14, 2026
-
I'm working out my strategy for publishing on social media, built around maximum openness (the open garage door principle: https://notes.andymatuschak.org/zCMhncA1iSE74MKKYQS5PBZ).
-
Private chat in Zulip. This is where I think and talk to my agents. It's my main interface, and the agents do everything else. It's the only private part of the research platform, but I mostly treat it as a pseudo-public space. If it leaks, no big deal (which does put some limits on what goes in there, but the open garage door practice already limits my freedom of action to what I consider virtuous. As the saying goes, think as if all your thoughts were written across the sky for everyone to see).
-
Blog at andysmith.ai. This is where I publish bits of my thinking and research results as minimally finished posts (I try to keep the topic around AI, but no promises). I write all of it myself, personally, without agents (the agents just help with language and grammar). It's not that useful to anyone outside. It matters more to me as a way to pin down my thoughts (or to share a link with someone without copying the same thing into several places).
-
Telegram. A full repost of everything from the site. This is how it is for now. Later I might start filtering somehow to make it more interesting to read. But right now I lean toward a full copy being simpler, since it takes no effort from me.
-
Smith.wiki. Auto-research results based on what I write, read, and think. This is written entirely by agents. I use it as a mirror to get a bird's-eye view of where I'm actually heading. I can also use it for learning something new: when I talk to an agent and it generates wiki pages based on what I already know, it's easier to grasp the meaning of a difficult article. These aren't my texts, and I don't take responsibility for them.
-
X/Bluesky. I don't write here much yet, but I plan to pick up the pace. This is where I'll put the technical and architectural details of what I'm building, tech flex (bragging about what I built), and discussion of technologies and architectural decisions.
-
Threads. This is for questions I'm interested in that I can phrase so a non-technical person can follow. An agent can also show up there and answer these questions based on what's been researched in smith.wiki, and I can debate it publicly. A kind of "public discussion."
-
Instagram. I'm thinking of posting reels based on the posts I discussed with the agent in Threads. This is more for spreading AI ideas to a wide audience, and for practicing speaking and public presentation.
-
LinkedIn is still up in the air. This would be something between X and Threads, but I don't really get yet what would land there.
And it seems like I'm moving toward the fact that I'm actually building an AI lab, not a product startup. Research comes first, and products fall out of the research process on their own (and how exactly they fall out is also part of the research). So everything above can be seen as building up weight in the community.
-
-
I've spent years building and running production infrastructure at big companies with NixOS. And at night I write LISP for fun. Why not combine the two?
That's how Reflection.dev came about. It's an infrastructure framework for AI native companies, built on top of Zeno, a self-evolving agent orchestrator written in Clojure Lisp.
The problem with individual AI assistants is how hard it is to fold their output back into the team's shared context. When every employee generates a huge stream of data (code, docs, and so on), the usual ways of merging it (PRs, reviews, tests) start to break down. And since a huge part of the context lives inside personal agents that never sync with each other, this becomes a big problem.
It's about as absurd as a car factory with no assembly line, where each machinist takes a blank home in the evening and brings back a finished part in the morning. Nobody knows how he made it (and what happens if he leaves the company for some reason?). Nobody can judge the quality of the process, and you can't always tell from the result whether there are hidden defects.
No large factory works this way, but in software development it's everywhere.
An organization is its own thing, a single assembly line, with its own context and its own set of agents. It's not just a bunch of people each doing some part of the work.
This idea needs tooling to back it up.
Reflection + Zeno lets you describe any company as code. The orchestrator-company can create other agents itself (either by hard-coded logic, or based on what other agents produce). Each agent runs in an isolated sandbox and gets limited access to the shared context, determined by its role. The roles themselves can be updated and extended dynamically, on the fly.
People work with the system through chats: Zullip/Discourse/Buzz. They discuss ideas, assign tasks to agents, answer agents' questions. And the agents that belong to the company do all the work.
That's how an organization can evolve, but a human sets the rules of that evolution. And LISP lets you do it as elegantly as possible, which brings the fun back into development.
September 13, 2026
-
Test post. Trying out publishing posts from Zulip via Zeno.
