Andy Smith

Building scalable self-improving infra for AI agents.

August 13, 2026

  • I'm reading about Gwern's Guardian Angels (https://gwern.net/guardian-angel).

    The idea is that instead of using general LLMs (which are trained to be useful to everyone, hackers included), you train your own personal LLM.

    Instead of putting some bit of info about your personality into the context (which is never complete by definition), the idea is that the "persona" gets derived from the data.

    The data is an append-only log of text editing actions. The model predicts the next action and suggests it to you. The success test is: the model wrote the whole text from the first paragraph.

    Right now the author's idea is to digitize his data as a stream of actions. He uses Emacs, so he gets this out of the box. There's a bit more detail here: https://gwern.net/nenex.

    Training a separate model for each user (and retraining it on new data on top of that) is expensive, so you'd have to use a simple model. If the model struggles to predict, it can ask a more expensive general model, ask the user for feedback, and fine-tune on that feedback.

    This is a cool thing. As an Emacs fan I want it, and not just for text and code but for any kind of activity.

    And there's another conclusion here that I agree with. Identity is exactly the actions we take. It's worth collecting them now, so there's something to train on later.

    The format, and how you capture, safely store, and use these actions, matters a lot too.

  • The classic SaaS business model is this: the user hands their data over to a provider, and in exchange they get an easy onboarding.

    The risks are obvious. The user doesn't know how the provider handles their data, and can't do anything about the risk of leaks. The provider can also lock the account without warning, or just stop existing, so there's a risk of losing your data forever.

    The opposite model is open source, where the user sets up their own servers, installs the software they need, spends time maintaining it, and takes care of backups, updates, and security themselves.

    The first way wins on scale. People seem to not care about the risk of losing their data.

    I don't get it. Is security and owning your own stuff really of no interest to people, and nobody even thinks about it? Or is it a conscious decision, and the point is that running and maintaining your own open source cloud is too complicated and expensive, more than the data is worth?

    And how much sense does it make to offer people a service for setting up and maintaining their own stacks for AI agents and their environment?

August 12, 2026

  • I've used PostHog for years as an open source alternative to Google Analytics. Today I read the docs and realized it's a deeper tool that flips the whole approach to infrastructure monitoring.

    Traditionally the technical data about the state of a system (logs, metrics, traces) lives in a separate engineering system (ClickStack, Grafana LGTM, DataDog, etc). In that model, the system is the center of gravity. Monitoring answers questions like "is microservice X working correctly" and "is there a problem with the database".

    For an engineer, an error is a line in a log and a ticket in an issue tracker. For the business it's lost money and lost customer loyalty.

    PostHog builds observability around the customer instead. Every event (including cases where something is broken), every log entry, every trace gets tied to a user session and stored right next to the business metrics.

    That inverts how you look at infrastructure, analytics, and monitoring. The system exists to serve the customer, not the other way around. That shift takes some getting used to. For years we've measured service uptime, not the customer's path.

    Now the cost of an incident is obvious. A broken payment button isn't a line in a log, it's the sessions that dropped off at checkout, and you see it in the same interface where you look at conversion.

    You shouldn't treat PostHog as a replacement for your technical observability stack. Some internal events aren't directly tied to user actions, so PostHog is an awkward place to look at them. But alongside the technical stack? Absolutely.

    image

August 11, 2026

  • What to do about tags on a blog.

    The point of tags is to highlight the topics I keep coming back to.

    One problem with tags is that they're unstable. They depend on how deep you've gone into a topic.

    At first, when I'm just starting to dig into a new topic, say AI, every post gets the same high-level tag: AI. Then as things get more specific, LLM shows up, then ollama, then mlx, and you can keep going deeper forever.

    But if you start with detailed tags from the beginning, you end up with a huge pile of empty tags, each one marking a single post. That's not useful. It's easier to just use search.

    For my blog I set up automatic tag generation for each post, based on the text.

    Each run starts from scratch, with no hint from a global tag dictionary. That keeps it from drifting toward the tags I used in my earliest posts.

    The problem this creates is duplicate tags. When the same thing is written slightly differently, like agent and ai-agent and ai-agents all meaning the same thing. With independent runs and no hints, you can't fully avoid that.

    For now I've left it as is. Later I'll add a periodic merge of tags, or some extra tooling to classify them and pull out common themes.

  • I needed to transcribe a call between two people.

    I looked for something ready-made and didn't find anything that fit. I wanted it to just work in one click on Apple MLX, split up who said what, and run LLM post-processing on top (to strip out the mumbling and filler words).

    I decided to throw together a small Python script, and as usual it grew into half a day of debugging (which is why I don't like vibe-coding my own tools).

    But I did get to do a quick bit of research on the state of speech recognition.

    For the model I tried whisper-large-v3-turbo, but it works badly, it actually makes mistakes. Right now I'm trying whisper-large-v3, and if it's no better I'll switch to parakeet, which from what I can tell works really well.

    For diarization (figuring out who said which line) we tried a bunch of options:

    • Sortformer didn't fit in memory on anything over an hour, and once we chunked it, it started glitching and counted 4 speakers instead of two
    • sherpa either split it into too many speakers, or collapsed and gave all the text to one, basically it glitched too
    • we tried a homegrown solution: cut on pauses and compare by embeddings, bad. either our voices turned out to be the same, or the embedding algorithm is bad. or maybe the voice doesn't factor into the embedding and the meaning was pulling the coordinates around
    • pyannotate seemed to work best of all.

    But there's a problem with the mlx + pyannotate stack, they need different pytorch versions, so they don't run in the same venv, I had to split them.

    On top of that I bolted on the LLM cleanup, and got a more or less readable result.

    Now I'll try to clean it all up, polish it, and open-source it.

    While coding it, I forgot what I needed it for.

  • I figured I'd look into how speech-to-text is doing these days.

    Locally, I use Handy for input.

    image

    I tried comparing a few tools, like openwhispr and wisprflow. Handy isn't the most feature-rich of the bunch, but the fact that it's free, open source, and runs locally wins me over.

    Compared to wisprflow, for example, it's instant. Because the request doesn't go anywhere, it's processed locally.

    Compared to openwhispr I didn't notice much of a difference, except that it's free.

    At first I thought the recognition model alone wasn't enough and you'd need LLM post-processing, but it turns out it's totally enough. LLM post-processing adds a lot to the processing time, and the quality only goes up by about five percent.

    Worth noting though, I don't need any text-editing features, meaning I need the text typed out word for word. If you want a transform option, say when you dictate something and the result gets run through some prompt, then Handy obviously won't cut it.

    But for my cases it's more than enough.

  • In the age of AI, it's really tempting to automate blogging, build content factories, all of that.

    But you have to split this by what you're writing for.

    Marketing copy you can and should write with an LLM. The LLM makes it better, tunes it to the right audience, checks for mistakes, and so on.

    But the stuff you write for yourself, you should write yourself.

    It doesn't matter how deeply the work gets automated, how many teams of AI agents I orchestrate, what percentage of the work is automated, or how deep the agents are in my work.

    The thoughts in my head stay. And they need to be made explicit (written down, and ideally published).

    Yes, the scale of the thinking goes up. I used to think like a developer. Now I think in bigger categories: like a manager of a team of agents, like a founder of a company of agents.

    The categories in my head are different, but that doesn't mean pulling those thoughts out and putting them on the page can be automated.

    So my blog stays written from my head. Sure, the thoughts will sometimes be naive, sometimes wrong, sometimes half-polished. But they're mine.

    That's how I split my own thoughts from the derivatives.

    And the derivatives, the ones prepped and made interesting for specific audiences, those can go out on social media. That content factory is the one I'll have to build for myself (and maybe I'll turn it into a product).

  • I spent a few hours today and yesterday working out a strategy for social media and being public.

    I want social media to be a nice storefront, showing only the good trail. But what I write is raw thoughts, from the angles I actually care about.

    I ended up landing on a split between "author-based" writing and "reader-based" writing.

    Reader-based writing is aimed at the reader. It solves some problem the reader has, it's interesting (otherwise the reader doesn't subscribe).

    But for me, as the author, the raw thoughts are the interesting part.

    The first solution I came up with was discobrain, a private notebook built on Discourse. Raw thoughts go in there, and I pull the "quality" posts out of it.

    But then I remembered https://simonwillison.net/ and got inspired to publish raw thoughts on my site instead.

    This means I have to keep an eye on some minimum level of quality, but it also lets me publish less polished stuff (nobody sees it unless they go looking for it).

    And there's still value in it. Maybe at some point I'll need to link back to some raw thought or idea.

    And on top of that, I don't have to filter by interest, because there are no subscriptions on the site. Nobody gets bothered, nobody unsubscribes, and no algorithm breaks if I post something irrelevant.

May 25, 2026

  • MCP is now used everywhere, and some products exist only as MCP servers. Some of these MCPs contain dozens or even hundreds of tools, and that creates a real problem when working with them.

    The problem

    The manifest of such an MCP is shipped in every LLM call, which means token usage grows proportionally with the size of the tool catalogue. The cost of running an agent system that depends on these MCPs grows accordingly. Prompt caching lets you reduce that cost (see IBM's overview and Anthropic's documentation), but it doesn't solve the problem of context window pollution.

    There's also a defocus effect: a weaker LLM can struggle to choose the right tool from a long list, and that affects the quality of the result. The Berkeley Function Calling Leaderboard (BFCL) measures function-calling quality directly and shows that smaller models visibly degrade. ToolLLM frames the same regime as a learning problem: how to teach an LLM to work with 16K+ APIs.

    The root issue is that the full information about every tool is sent on every request — while in practice only one or two tools from the entire list will actually be used.

    An optimization

    This can be optimized. Suppose we have an MCP for working with notes, and we want to implement two basic tools: create(text) and find(keywords). Instead of implementing them as separate MCP tools, we can expose a single one: notes_mcp_call(method, params). Then we'd invoke them as call('create', ['Hello, World']) and call('find', ['notes about emacs']).

    Effectively, this is untyped RPC dispatch on top of MCP, whereas the conventional approach is to expose each method as its own typed tool. Naturally, this isn't a silver bullet — choosing one approach over the other is a real architectural decision for the MCP developer. The proposed approach pays off most when the number of tools is genuinely large. And for dynamic MCPs, where the set of methods isn't known ahead of time, it's arguably the only viable option.

    The discovery problem

    With this approach, a discovery problem appears immediately. The LLM has to learn somehow what methods this MCP exposes — but at the same time it shouldn't receive the full list of tools all at once.

    So we also need a service method, help. It can be implemented as call('help', ['create']), or as a separate MCP tool.

    Regardless of how help is implemented, it can operate in several modes.

    The first mode is keyword search (or semantic search). The LLM asks the MCP something like "how do I create a new note?" and gets back a list of tools with their descriptions, parameter lists, and result descriptions. This is very easy to implement, but it requires maintaining semantic search infrastructure, including embedding the incoming queries. There's a bigger issue, though: building a complete map of available methods is hard for the LLM, because it doesn't know all of the MCP's capabilities — and so the discovery goal isn't actually reached. The model has to already know what it needs from the MCP.

    The second mode is FSM-style discovery: help ships wiki-like documentation with cross-links. help without parameters returns a general overview and links to other pages. The LLM reads the wiki sequentially and assembles information about all the pages. This mode can be combined with search; it enables real discovery, but it requires extra effort from the MCP developer to maintain that documentation. As a side note: nothing stops you from implementing just a single FSM state with the full list of tools — in that case the mechanism behaves very close to the default MCP behaviour.

    The final design

    So the MCP manifest ends up containing two tools: help and call. The description of call should include the call format, instructions on how to use help for discovery, and optionally a description of the most frequently used methods — so the LLM doesn't have to go through help for every little thing.

    What already exists

    Before building this ourselves, let's look at what already exists in the industry.

    Anthropic implements dynamic tool loading in its own products through tool_search, but this doesn't work in other vendors, so it can't be used as a universal pattern.

    There's a draft standard proposal, SEP-1821, which extends the MCP standard with keyword search. This is partly what I need, but it doesn't enable flexible FSM-style documentation.

    Speakeasy implements a very similar pattern in their tool Gram (see their blog post and the documentation).

January 30, 2026

  • I've been reading Anthropic's research on the problems with AI assistance when acquiring new skills (arXiv paper).

    The hypothesis: AI accelerates your work where you're already an expert but hinders you from becoming an expert in something new. If you fully delegate tasks to AI, you start to get dumber and eventually become obsolete.

    At first glance, this contradicts my post on autonomy with acceptable quality, but it doesn't. In that post, I was discussing autonomous task completion, not learning something new. I set the tasks myself, which means I already have some understanding of what needs to be done and how to evaluate the result.

    But what do you do when you lack that understanding? How should AI help you learn when it always wants to do everything itself?

    The research suggests these approaches:

    1. Discuss conceptual options but write the code yourself
    2. Generate code and ask for explanations

    The key point: you need to put in the effort, to think. That's when learning happens. If you just mindlessly delegate task solutions to AI, you won't learn anything.

    My Approach

    I usually take a different path when learning something new or building fundamentally new (for me) systems. These are tasks I can't yet fully delegate to autonomous AI:

    1. Problem formulation: Define what needs to be done and how to evaluate the result
    2. Research: Gather all possible solutions, existing technologies, tools, and practices. Compare and choose
    3. Architectural decomposition: Together with AI, I build an understanding of how the finished system will work. It's crucial for me to actually understand what the result will be because I'm responsible for the decision made and implemented. I'll need to review and accept the result. This is, in my view, the key difference from mindless vibe coding. Without this understanding, there's no way to take responsibility for the outcome, and the result may or may not work out
    4. Documentation/tests preparation based on the discussion above. AI formulates, I carefully reread these documents and make many changes. This is where my understanding and AI's understanding synchronize
    5. Writing code: Here I mostly trust the AI. I only review the most critical parts, but I believe verification should be automated through tests (including architectural ones) and possibly other methods (formal verification)
    6. Verification: Automated through tests
    7. Acceptance: Still manual, but this needs to be automated too
    8. Debugging: Automated by AI, but it's important to engage rather than just copy-paste errors. I always require descriptions of error causes and read them to understand what actually happened

    After going through this cycle for the first time, I can extract this class of tasks into an autonomous solver that will do all of this next time without me.

    Understanding is the key in this process. Without understanding, you have problems.

    Is Shallow Understanding Really a Problem?

    Maybe it's becoming the new normal? Maybe I'm the odd one? After all, Anthropic didn't put an abacus in the header image for nothing. Who knows how to use an abacus in 2026? Who even remembers what it is?

    Building AI systems to work reliably is also a skill you can learn (including experimentally), putting in effort and making mistakes. So this also needs to be done, which is exactly what I'm doing and describing in this blog.

    But as usual, the truth is somewhere in the middle. Both matter.

  • When actively using Claude Code in manual mode, I always apply the same technique. I keep sessions as short as possible. I set a task, get the result, close the chat. If I realize I need to return to the task to clarify something, I do so via claude -r.

    This workflow creates a need to save information between contexts. At the end of each session, I ask Claude to summarize and format a short summary of the session outcomes, which I then publish to the issue tracker for that task. I also ask it to update documentation, tests, architecture decisions, and CLAUDE.md to keep all descriptions synchronized.

    I recently heard the opposite recommendation: do everything in one session, never close it. The guy complained that his tokens were flying away catastrophically fast, but when I suggested keeping sessions short, he said someone had recommended doing everything in one session.

    In that case, the context contains not just what you need (the things you deliberately identified as important and placed in CLAUDE.md). It contains absolutely everything that was discussed, and half of it gets lost anyway. This is clearly an inefficient way to interact, which that guy discovered in his own wallet, but he refuses to believe the obvious because faith in an authority's words turned out to be stronger.

January 29, 2026

  • The primary metric I use to evaluate my autonomous AI systems: autonomy of work with acceptable quality of result. This seems obvious from the name "autonomous," but until you articulate it explicitly, it's not.

    I see people around me setting up multiple monitors to watch several parallel Claude Code sessions simultaneously, constantly tweaking and running from one to another.

    I believe this approach is fundamentally wrong. Context switching in the human mind is an expensive operation. Very expensive. Frequent task switching is exhausting, regardless of what anyone thinks or says. There's research on this: The Cost of Interrupted Work, Executive Control of Cognitive Processes, Brief Interruptions Spawn Errors.

    So my job as an architect of this class of solutions is not to "do as much as possible with AI," and not simply to "efficiently burn tokens," as I thought before (see also From Solo Sessions to Agent Orchestras). It's specifically to ensure autonomy with acceptable quality.

    That means I need to ensure predictable and repeatable results with minimal effort on my part.

    I don't measure how many tasks I did with AI. I measure how many tasks AI did without my help. Of course, it's not entirely accurate to say "without my help" since I built the system that enables it to work effectively and autonomously. But that's exactly the point.

    This is about the extent to which my solutions are AGI.

  • My hypothesis is simple. The recipe for success is to offer something simple and familiar (lowering the barrier to entry) and add value on top of it. Products that work on interfaces users already understand take off. Products that completely change how people work don't.

    Consider these examples.

    Cursor is VS Code, familiar to everyone. You don't need to radically change how you work. Just tweak your workflow a bit.

    Claude Code is a terminal that every developer already knows how to use. AI-powered development (the value) is built on top of it. Developers don't need to get used to a new interface. The barrier to entry is minimal.

    Such products don't require users to invest significant time changing their process. That's the key. This is the quality you must preserve when building your own product.

    And there are thousands of examples of products that force users to work differently with their calendar, their code, to change all their daily habits. They don't take off and never will.

    If using my product requires the user to perform many actions and learn new things, meaning they must take a very wide step from their current state (not using my product) to the target state (spent time, uses the product, receives value), then even if the benefit is obvious to them, they'll likely give up without even trying.

    However, once a user takes that first step, they'll take subsequent steps with much more enthusiasm and commitment (because they already feel the value).

    This is what you should always keep in mind. Simple beats complex. And something simple is better than nothing at all (see Done Today Beats Perfect Never).

    Many small steps beat one big leap.

January 28, 2026

  • Yesterday all channels were buzzing about Clawdbot. I decided to install it and give it a try.

    I liked the concept. Following my thinking from Product Over Technology, I always consider the product concept and how to sell it first. The execution is not great, but that's not what matters. What matters is that it's generating hype (meaning it sells) and it works well enough.

    What I fundamentally didn't like was the lack of a platform approach. No separation of important and unimportant. No core versus everything else. This led to having to think about everything during installation. Network configs (Tailscale), provider settings (though you could reuse what's already in ~/.claude/), plugin sets. I have no idea which plugins to enable right now. I haven't even decided what I need this for. I want to try it quickly and move on. In some places, I noticed NIH syndrome (Not Invented Here), with vibe-coded solutions instead of existing tools.

    Because of this, the whole thing looks like a super over-engineered solution tailored to one specific person (the developer), ignoring the fact that different people need different things.

    What I Would Do Differently

    I would extract the core that absolutely everyone needs. Which I actually did. My Capsules are exactly about this. Then I would let the bot grow and develop itself.

    This actually looks like a solution to my problem of agent interaction in Capsules. Instead of some unified supervisor orchestrating everything, I can let agents self-develop and interact with each other.

    To demonstrate how Capsules work, I could launch my own personal assistant bot (similar to Clawdbot) running on Capsules with my own vision. In my vision, the bot-assistant doesn't create other systems itself. Instead, it provides communication between the user (the botlord) and other agents. And possibly other people through their assistants.

    I see the core as super minimal. Just a chat through which you can give commands for self-improvement. Through this chat, you configure the bot itself. Self-development. The goal is to get a working companion in seconds and then tune its capabilities over time, including various access interfaces.

    This resonates with my idea to use Nix for describing Capsules. In this case, the bot can literally write itself.

    My hypothesis is confirmed: Clawdbot takes off because it uses standard familiar interfaces (Telegram, Discord) and adds value on top of them. The interface matters. I expanded on this in Small Steps Beat Big Leaps.

January 27, 2026

  • I've always tried to see myself as a company. Even as an employee, I viewed my employer as a client or partner. The problem is, I did it poorly. I only recently realized that I think in processes rather than product-outcomes, and that needs to change (see: Product Over Technology). But the core idea remains: I am a single independent economic unit that joins forces with other independent units to achieve shared goals. This partnership is mutually beneficial and voluntary.

    With AI, this approach intensifies. The boundary between individual and company is dissolving. This video proposes viewing yourself as a complex AI corporation, which creates new challenges people haven't faced before: management, finance, security, oversight, legal, and more.

    The proposal is to become a manager. Not a seagull-manager, department head, or paper-pusher, but a true CEO: a leader and visionary. Apply classical management models to an organization of your AI agents. You'll need to build organizational structure, design data processing pipelines, delegate tasks and decisions, and take personal responsibility for your employees' actions. This means not giving them too much authority, or better yet, running them in isolated capsules. There's no other way. Everyone must learn to build systems, not just do things. Push yourself left along the value chain!

    Now is the time to learn management.

    Though I see a challenge here. AI agents aren't people in the traditional sense. You can't punish or reward an AI agent. Classical models will need adaptation. But the core principles can probably still apply.

  • I've spent my entire life working with technology. Studying it, building it, growing it. When evaluating any project, I instinctively reach for the technology lens first. How well is the code written? Does the architecture allow for future growth? What frameworks are being used?

    When developing my own projects, I always focused on technology first. I design elegant architecture, set up auto-deployment to Kubernetes, ensure data security, scalability, and disaster recovery. I always have monitoring in place.

    But I'm missing the main thing: sales. Because my project isn't about business. It's about technology. I'm just a kid who never finished playing with blocks or construction sets. I find it interesting to build a system not to make money, but to say: "Look at this sandcastle I made, isn't it beautiful?" and then walk away to start the next "project."

    This realization was as unexpected for me as it seems obvious in hindsight.

    I think this is a very common problem among engineers who spent years working as employees. "Business isn't my thing, there are other people for that," they think. But wait. Isn't your life your own business? Are you really willing to hand over the right to manage your life, to take responsibility for the outcome, to someone else? Because when those people make mistakes, you're the one who suffers, not them.

    It's time to take responsibility. It's time to decide that we're doing business first and building second. What matters is why we're doing this; the how is secondary. It's time to focus on the product, on markets, on economics, psychology, sociology, and other aspects of business. Formulate hypotheses and test them instead of waiting for permission from someone. What's the product concept? Who will pay money, and for what? Why would they pay for our product instead of another?

    Everything else? A team of AI agents will handle that.

January 22, 2026

  • My Evolution of Working with AI Tools. The Capsule Concept.

    I use LLMs and coding agents extensively in my work. This post contains a brief history of my observations and hypotheses.

    The first decision I made: never run agents locally. This is a rule for me. There are two reasons for this. First, I'm afraid the agent might accidentally do something unacceptable on my behalf and with my permissions. For example, delete my home directory or a client company's database (protection against such actions may either be absent or fail). Second, the agent could steal secrets from my device (keys, wallets, passwords), especially if some third-party MCP or skill is used that might contain a prompt injection instructing it to send the entire contents of ~/.ssh to an attacker's remote server.

    Acting according to this decision, I started running all development on a remote machine. But I immediately encountered the next problem. For different types of tasks I need different settings, different sets of MCPs, skills, different environments. I want something like "profiles" (https://github.com/anthropics/claude-code/issues/7075), but more capable. I want a profile to include not just the Claude Code state but the entire environment: all necessary MCPs, all keys, tokens, passwords required for the agent to work, repositories.

    This leads us to a new concept that I call Capsule. I'll provide a more complete architectural description and decisions about the capsule's internal structure later, but for now let's consider external interaction with a Capsule as a black box.

    From a conceptual standpoint, the following aspects should be considered:

    1. Capsule Template. A textual description (config) listing what should be inside. Having a text format here is critical, as it will allow using git and LLMs to create and manage capsules.
    2. Capsule Instance. Creating a Capsule from a template. Multiple Capsules can be launched from a single template.
    3. Command Interface. Giving commands for execution. The agent inside the Capsule needs some way to receive instructions from outside about what to do (not mandatory, some agents can go out into the world and request instructions themselves), as well as provide statistics about its work. This will allow people to interact with the agent in a capsule and agents to interact with each other.
    4. Other Interfaces. MCPs can access anything, but strictly according to the rules described in the template.

    Since we're talking about wanting to describe the entire environment, Nix/NixOS fits this concept very well. Most likely the template will be a Nix Flake, the instance will be a virtual machine deployed by Nix, the command interface will be determined by the agent, and other interfaces by MCP servers and other tooling in the environment. But these are still open questions.

    The next question is agent selection. Currently I use Claude Code (and only it). But I need a command interface as an API and preferably a web interface, so I don't have to SSH into the virtual machine. In this case I'll probably look towards opencode, since it supports API out of the box, and in my view the quality of work for all open-source agents will tend towards the same value (since the code is open, successful solutions will spread to all tools, and unsuccessful ones will die out). Therefore there won't be a big difference between choosing opencode or Claude Code.

    The immediate plan is to develop 1-2 capsule templates for different tasks and start using them. After that, I can think about how to run multiple agents in parallel.

January 21, 2026

  • The tactic is simple. Figure out how to consume all available compute, then find ways to increase that availability.

    I have a Claude Code subscription and other AI tools. My goal is to use them as efficiently as possible. This means pushing token usage closer to 100% of the available limit while maintaining acceptable quality. Solving real problems, not running expensive LLMs in circles.

    Measuring Utilization

    For objective assessment, I need to periodically collect usage and limits statistics. Perhaps every minute or every ten minutes. Store it in a database, build graphs, and devise methods to minimize the area under the (limits - usage) curve across all three parameters: current session, weekly totals, and Sonnet-specific allocations.

    How to objectively evaluate effectiveness? I haven't fully figured this out yet. It will likely be an LLM bot that reads all my sessions, looks for anomalies, and suggests improvements.

    Scaling Options: From Simple to Complex

    Here's how I see the path to full utilization:

    1. Parallel Manual Sessions

    The simplest approach. I work on 2-3 tasks in side-by-side Claude Code sessions, switching my attention between them. Low overhead, immediate results.

    2. Agent Orchestra

    Design an AI environment where multiple agents work in parallel, communicate, negotiate, and I simply observe. This requires upfront architecture work but multiplies throughput.

    3. Overnight Research Tasks

    Formulate tasks for long-running research and leave them running overnight with a defined token budget. Wake up to results. This captures hours that would otherwise be wasted.

    4. Periodic and Event-Driven Tasks

    Assign the AI recurring jobs. Collecting daily email summaries at 4 AM, responding to certain events. These tasks utilize nighttime hours when I physically cannot participate in options 1 and 2.

    5. Fully Autonomous Agent Teams

    The end goal. I set a task, agents self-organize into teams, and solve problems while planning their token expenditure based on usage statistics. They schedule research and tests for nighttime (compute-heavy, LLM-intensive, no human needed) and communication tasks for daytime.

    The Dual Scaling Problem

    This breaks down into two challenges. Vertical scaling: teach agents to work autonomously for longer periods. Horizontal scaling: teach agents to collaborate effectively as teams.

    I'm confident many teams are working on this. I'm happy to contribute what I can.

    The Resource Allocation Problem

    One thing I haven't solved: how to distribute usage between my manual work sessions and automated agents. If agents consume all available tokens, but I need to do something myself, what then?

    This looks like a case for applying management and organizational planning practices to compute resources. Budget allocation. Priority queues. Reserved capacity for human override. Essentially, accounting for AI agents.

    Solving these problems would mean the scaling challenge is addressed. From there, it's infinite incremental improvement.