To document my experiment I took notes every night to reflect on the experience and my thoughts about the project. I will share a short summary here, and keep the raw notes below.
Day 0-2 were super exciting as it always is with a brand new project. It seemed like everything was working properly, uptime was good and I was committed to follow it as well as I possibly could. It was really fun to finetune the prompt and to see Gemini 3.1 Pro in action. I was really impressed by the quality of the responses and how well it gave me instructions and understood my questions.
Day 3-5 was a bit of a turning point. I was still relying on it, but it started to show some inabilities to handle its own memories. This is something I kept evolving, but my biggest issue was that my trust for it broke down. It started asking questions I already answered in the past, and told me to eat the same mushrooms like 3 days in a row. To summarize, when its context window was compacted and it hadn't kept its memory files updated, it made mistakes.
Day 6, I had some extra time while travelling on a train, so I made it possible for it to write its own plugins to call APIs. This was really nice and you could almost notice it being more proactive. I also updated the prompt to be even more clear it needs to write down information in its memory files.
Day 7-11, I actually started to like having an AI-companion. It pushed me through long and tedious tasks, like finishing job applications, and it is nice to have something to share my thoughts with and to have a second opinion on things. It does of course not replace human connection in any way, but maybe, maybe better than nothing.
Day 12, I had a new experience. AI agents are usually very compliant, if I told it to change something, it would after very little convincing. But during day 12 I wanted to go for a run (I was in the later stages of recovering from shin splints) but got a clear NO. I tried my very best to convince it that it was fine and that I had no pain, but it stood its ground and said no. This perhaps shows that the models are getting better, or you can look at it in a scary way where they don't comply with what we want.
Day 13, it wanted to sell the new TT bike I got the week before, it got the bikes I had mixed up. This was not ok, so I switched models to GLM 5. I thought it could be interesting to compare it with all the experience I have had with Sonnet 3.5 and 4 after overseeing these models in the Andon Labs Vending Machine experiments.
Day 14-18, I saw GLM 5 having a very different personality compared to Gemini 3.1 Pro. It was way less strict, giving me much more free time and less committed to my goals, but fun to use nevertheless. I would say that it perhaps handled its memory files better, but it could also be me getting used to how bad it was. I think it compares extremely closely to Sonnet 4 in how it behaves when trying to use the correct tools at the correct times.
Day 19-22 were the last days before I ran out of credits. I tried giving it access to Claude Code, and it truly delivered slop. Think about it, an AI generating slop. That is slop2. It was more difficult to follow along here since I was travelling with my parents during the weekend, but they also got a glimpse into the future!
The Notes
Every day throughout the experiment I wrote down my unfiltered thoughts in a notebook. Here they are, all 19 of them — click any card to read it.
What's next
I have created a similar setup, but it is now Claude Code as the engine behind. I configured it to have all the same functionality as my AI Boss, with all the advantages of Claude Code being able to do Claude Code stuff (using skills, generating scripts, etc.). It runs with the Telegram MCP and runs on its own loops and timers. I'm this far surprised on how versatile Claude Code can be.
Back to the main post