MS-004 · Software
Claude Code
Overview
I have been working with Claude Code since the middle of July 2026. It is a command-line coding agent made by Anthropic, and I run it inside Cursor on a Windows laptop. Everything below this line was built with it: this portfolio, a personal dashboard, and a live property management application for a real business.
Most of what I have spent time on is not the code itself. It is the setup around it. Claude is the head and other models are the hands. The main model makes the decisions, judges what comes back, and talks to me. Long or repetitive writing goes to free models from other providers. Multi-step work that needs tools, like searching a codebase or reviewing a change across several files, goes to smaller sub-agents. I decide what goes where before the work starts.
What goes where is not a guess any more. At the end of August I measured where my own sessions were actually spending, and every routing rule I use now came out of that measurement rather than out of what felt fast.
Why route anything at all
A subscription comes with a usage budget, and the strongest model burns through it fastest. If it spends that budget writing boilerplate, there is less of it left for the part that actually needs judgment. So the expensive model does the deciding and the checking, and the cheap models do the volume.
The router
I built a local router that puts several free providers behind one command, so a prompt can go to whichever model suits the task without me changing tools. It sits in front of nine free provider accounts, among them Groq, Google Gemini, Mistral, NVIDIA, Cohere, Cloudflare, Hugging Face, and OpenRouter, and between them there are several hundred models it can reach. That part was built on August 28 and 29, 2026, after two months of doing it by hand.
Which of those models actually get used is not fixed. Every call is logged, and a task goes to the models that have been answering, not to the ones with the best reputation. When one starts failing it comes out of the rotation and something else takes its place.
It can also send one prompt to several models at once. Because they answer in parallel, getting five answers costs about what the slowest single answer costs. The models are deliberately picked from different companies. Several versions of the same model agree with each other and teach you nothing.
Picking between the answers
Reading five answers costs more than it saves, so I judge them mechanically first: count the lines, run the test, check the constraint, compare them against each other. Only the one that survives gets read properly. The rest never enter the conversation at all.
Where it stops working
Free models are good at producing text and bad at tool loops. Given a job that requires several steps with real tools, they stall or invent the tool call instead of making it, so that work stays with the Anthropic sub-agents even though they are not free.
They also invent specifics with total confidence. Asked once for résumé bullets with no figures supplied, a free model returned two percentages that were never in the prompt and had no source anywhere. Nothing about the output looked wrong. So every number, date, and name that comes back gets checked against something real before it is used, and this page was written the same way.
Measuring instead of guessing
For the first two months I tuned all of this on instinct. Something felt slow, so I changed it, and I had no way of telling whether the change had helped. At the end of August I stopped guessing and parsed my own session logs instead: every call I had made since July, and what each one actually cost me.
The answer was not the one I expected. The thing I had assumed was broken turned out to be working correctly, and nearly ninety percent of the cost was coming from two completely ordinary habits: opening whole files to answer small questions about them, and taking screenshots of pages in order to read text off them. Both had a cheaper tool sitting right beside them the whole time. Searching a file instead of opening it is roughly thirty times cheaper, and reading a page as text instead of as a picture is cheaper by more than that again.
What the measurement changed
A finding only matters if it changes what happens by default, so these became standing rules in the instructions file that gets read at the start of every session: search a file rather than opening it, read a page as text rather than as a picture, start a clean session when the subject changes, and send long output to disk and review it in pieces instead of reading the whole thing back.
The instructions file itself got cut roughly in half in the same pass. Everything that was reference rather than rule moved into a second file that is only opened when it is needed. That file is re-read on every single call, so its length is not a one-time cost. It is a small tax charged thousands of times a session.
Making the cheap path the only path
The change that did the most was the one I do not have to remember. There is a setting that caps how much any single tool call is allowed to hand back, and lowering it makes the expensive habit impossible rather than merely discouraged. A rule you have to remember is weaker than a limit that will not let you.
The same pass cleared out what was not being used. I checked which add-ons had ever actually been called across every session on disk and switched off the nine that had not. Several of them had never been used once.
The three sessions I spent blaming the wrong thing
Delegation stopped working at the end of August. Prompts hung and came back with nothing, and across three separate sessions I concluded the router was dead and worked around it.
It was not dead. It had been routing every call correctly the whole time, and the log on my own disk said so. I had never opened it. The default route pointed at a model that was rejecting the requests for being too large, the fallback behind it was timing out, and each attempt was allowed five minutes with no limit on how many attempts it made. A two-word prompt was spending a minute and a half failing four times over before it gave up.
The fix was to route off the log instead of off reputation: pick by the success rate a model has actually shown, cut the per-attempt timeout to well under a minute, and stop after two providers fail instead of working down the whole list. The same two-word prompt now answers in under two seconds. I also built two small tools to read that log, one that reports what the free models have actually done and one that finds capacity I have never tried. Neither of them calls a model. They read files I already had.
That is the part I would keep if I kept nothing else. The evidence sat on my own machine for three sessions, and the problem was that I had not read it.
What it costs and what it does not
The free models have handled well over a hundred thousand tokens of work so far without touching the subscription. That only holds because of a few rules I do not break:
- Everything that comes back gets verified, hardest on the numbers
- Nothing gets delegated if writing the prompt and checking the answer costs more than just doing the work
- Anything that needs tools stays with the paid sub-agents, because free models are bad at tool loops
- More agents is not delegation. Extra Anthropic sub-agents spend from the same budget; the free models are the only part that actually saves it
Every answer that used them ends with a short ledger of who did what, so the saving is something I can check rather than something I assume.
matthewsabag.com
The site you are reading. I started it on July 17, 2026 and it has stayed active since, 60 commits so far, 8 of them in July and 52 in August. It is plain HTML, CSS and JavaScript with no framework and no build step, plus a couple of serverless functions for the contact form and the now-playing widget. It deploys to Vercel.
It is also the project I learned web development on. The 3D model viewers, the PDF reader, and the video embeds are all here because I wanted to know how they worked, and none of them went in cleanly the first time.
Visit matthewsabag.com Live site
What it took
The first version went up and I redesigned the whole thing the next day. In August I reworked the background and the home page layout again. Neither rewrite was planned. Both happened because the first attempt only looked wrong once it was actually on a screen.
The 3D viewers took three separate rounds. The first time the module import map was wrong, so the 3D library never loaded at all and nothing appeared. Once that was fixed the models loaded but faced the wrong way. The third round was framing and camera angle, getting each model to sit in its frame at a size worth looking at.
The PDFs started out opening in the browser's own viewer, which works but looks like a different website every time. I ended up rendering the pages myself so the documents match everything around them.
The pattern is the same in all three. The thing I thought was one change turned out to be three, and I only found the second and third by shipping the first and looking at it.
Daily Dash
A personal dashboard I started in July 2026. It pulls the things I actually check every day into one place instead of five apps: lists, calendars, a fitness ledger, and a daily recap. Next.js and React on the front end, Postgres through Supabase behind it, deployed on Vercel. It runs on a preview URL while I build it, and it is still my own daily driver rather than something open to the public.
The two walkthroughs below were recorded eleven days apart. Putting them next to each other is the clearest picture of how the project actually moves, because almost nothing on the screen in the first one survived unchanged into the second.
What it does
Lists. A board of cards for to-dos, groceries, and school assignments. Items carry a priority and a due date, and they can repeat on chosen days of the week. A list can be shared with someone through a link.
Calendars. A month and week view that merges everything. It pulls my UF Canvas assignment feed and my Google Calendar automatically, so coursework shows up without me typing it in.
Fitness. A running ledger built on a credit system. Runs, rides, gym sessions, sports, and push-ups each earn credits, and the balance tells me how far ahead or behind I am.
Reminders. A Telegram bot sends a daily recap and nudges me when something is coming due, so I get the dashboard on my phone without opening it.
What changed between the two videos
In July the whole app was one screen: a board of draggable cards holding lists, notes, and an overdue panel. It worked, but everything lived in one place and everything looked the same for everybody.
By August it had been pulled apart into tabs, each with its own home. A summary dashboard replaced the board as the landing page. Calendars arrived and started pulling in coursework on their own. Fitness arrived with the credit ledger. Finances arrived with expense tracking and live market prices on a real chart.
The bigger difference is that the second one bends to the person using it. Settings grew into five tabs. Calendars can be renamed, recolored, and dragged into the order you want. Individual tabs can be switched off. The fitness credit rules are set per user rather than fixed. The July build had one opinion about how it should look and behave. The August build asks.
None of that was planned up front. Each piece got built because using the thing daily made the gap obvious.
What broke
Most of what I learned came from things breaking in ways I did not expect. A database read silently stopped returning rows past a limit I did not know existed, and the page went blank with no error to point at. Two calendars kept duplicating themselves overnight because of how the database treats empty values in a uniqueness rule. A sync running on every page load made the whole app feel slow, and the fix was to stop doing work the user never asked for.
The pattern across all of them is the same. The symptom and the cause were nowhere near each other, and guessing wasted more time than reading did.
August 2, 2026 · current
July 22, 2026
AtidRealty
A live property management application for a real business, and the only one of the three I did not start. I took over an existing codebase on July 19, 2026 and have been maintaining and extending it since, 69 commits so far. Node and Express on the server, Vite on the front end, Drizzle for database migrations, Radix components, TypeScript throughout. It runs on Hostinger.
The work has been in two halves. I audited the code I inherited and tightened up how it handles data, and I built out the parts the business actually uses day to day: tenants, leases, documents, and the accounting side.
Working on something people are using
This is the project that changed how I work, because it is the one where a bad push is not a bad afternoon. Deploys are numbered and the count ran into the sixties, each one a batch of fixes going out to a site people were using while I worked on it.
Reading the code I inherited took longer than changing it. Several problems were not where the report said they were, and the fix that looked obvious would have left the real cause in place. On my own projects I can guess and try again. Here guessing is expensive, so I stopped doing it.
The deploy pipeline
For most of it I was pushing changes up by hand, which is fine until the day you are tired. So I built a pipeline: push to the repository, and the app gets built and shipped to the server on its own. It took about six attempts across two days to work, and every failure was somewhere other than the code. A path that did not expand the way I expected. A step that failed intermittently for its own reasons. A formatting mistake in the pipeline file itself. Connections timing out.
That was the useful lesson out of the whole project. On something real people depend on, how the change reaches them matters as much as the change.