↓Skip to main content
  1. Notes/

Zeroclaw thoughts

I’ve been getting more comfortable with Zeroclaw as my main headless agent for day-to-day work. It’s written in Rust, and the out-of-the-box Docker support is excellent, so it suits my Raspberry Pi deployment well. Matrix support is solid, although fiddly to set up as always, and I think that’s just an inherent limitation of the Matrix protocol and the demands of E2EE.

It’s currently paired with a local instance of oMLX to provide local inference for the agent, running Swift-Qwen3.8-Flash-Next-oQ4e-mtp which is the MLX version of the UkisAI Swift1.5-Qwen3.8-Flash-Next model. The Swift family of fine-tunes are designed to reduce the overthinking these models are prone to, along with other similar projects like ThinkingCap from BottleCap AI. On the M1 Studio Ultra, I’m getting around 30-35 tokens per second with this model.

The other improvement I’ve made to my local setup is to apply a patch which allows me to use my local isntance of oMLX to serve a speech-to-text model so the agent can transcribe voice notes from Matrix rooms. To support this I’ve made a small patch which allows you to specify a model name when submitting requests to http://192.168.1.X:8081/v1/audio/transcriptions.

The particular model I’m using is another from the Qwen team, mlx-community/Qwen3-ASR-1.7B-6bit, which was originally announced back in September 2025. Both happily run within the 128GB of RAM this machine has.

The next step I’m working on is to self-host an Obsidian sync server so I can use Obsidian as a shared note-taking tool. In chatting with incredibly well organised Alistair Tse recently, he showed off an extremely impressive curated setup which I am envious of!