Coding with an AI agent without the cloud: Ollama and MLX on a Mac

AI coding assistants send your code to someone else's servers and want a subscription or an API key. With an Apple Silicon Mac you can run a model locally: no cloud, no usage fees, and it works offline. Here is what you need and what to expect.

· 4 min read · Enrico Fanucchi

Short answer: you need a Mac with an Apple chip (M1 or later) and at least 16 GB of memory, Ollama to download and run models, and a model that fits your memory. A local chat only answers; to have code written, run and verified you need an agent such as Ollocode, which works with local models only.

Why local

  • Your code stays yours: nothing leaves the Mac. That matters for client projects, code under NDA, and anyone who simply prefers it.
  • No usage fees and no API keys to manage.
  • It works offline, on a train or on a poor connection.

The price is power: a model that fits in a laptop's memory is less capable than the largest cloud models, and slower. For many tasks it is enough, if it is used the right way.

Which model for which Mac

Memory is the main rule: the model has to fit entirely, together with the conversation's context. These are the models Ollocode uses as its agent depending on the Mac, measured on 14 real projects to build or fix (Go, Python, Node.js, TypeScript, Rust, web), passed only when the automated tests pass:

MacModelProjects passedNotes
16 GB (e.g. Mac mini M4)qwen3.5:9b7 of 14small and medium tasks
24 GBgpt-oss:20b5-6 of 14very fast, less consistent
32 GB or moreqwen3.8:27b-mlx13 of 14the most reliable, slower

One useful finding: code-specialised models built for completion, such as qwen2.5-coder 7B and 14B, proved unreliable as agents in these tests because they misuse tools. For inline completion in the editor, though, a tiny model like qwen2.5-coder:1.5b (1 GB) works great.

MLX is Apple's framework for running models on the Mac's chip: MLX builds of a model make better use of Apple hardware, and that is the build used on Macs with more memory.

Installing Ollama and a model

  1. Download Ollama from ollama.com, install it and open it.
  2. In Terminal, pull the model that suits your Mac and try it:
ollama pull qwen3.5:9b    # for a 16 GB Mac
ollama run qwen3.5:9b     # a chat in Terminal, to try it
ollama list               # downloaded models and their size

Now you have a model that answers. But a chat is not enough for coding: it hands you a piece of code and it is up to you to paste it, run it, read the error and ask again.

From chat to agent

An agent does that loop by itself: it reads the project, writes to the files, runs tests and programs, reads the errors and keeps fixing until the work actually works. With local models this matters even more, because they make more mistakes: an agent that checks every step catches and fixes its own errors.

Ollocode is a free macOS IDE built around that idea, with local models only (Ollama or MLX):

Ollocode: the agent writes the code, runs it and reads the result
Ollocode: the agent writes the code, runs it and reads the result
  • it doesn't settle: work is not done while tests fail or while the latest change has never been run;
  • it adapts to your Mac: it reads the chip and memory, picks the model and context measured for that class of Mac and, if the model is missing, offers to download it;
  • Plan mode: before touching any file it proposes the steps, and you adjust them until you are happy;
  • remote projects over SSH: it works directly on a Linux server, for example with PHP and MySQL;
  • built-in editor, diff, Git and terminal, with inline completion from a local model.
Ollocode picks the model based on the Mac's chip and memory
Ollocode picks the model based on the Mac's chip and memory

Safety: an agent that runs commands

An agent that runs programs needs guardrails. In Ollocode commands run in a macOS sandbox that only allows writing to the project folder, temporary folders and tool caches; a restore point is created before every turn that changes files; and system commands, such as installing software or using sudo, only run when you press Run. Reading the web for up-to-date documentation asks for permission too.

Every system command is shown in full and runs only after you confirm
Every system command is shown in full and runs only after you confirm

What to expect, honestly

With 16 GB a local agent handles small and medium tasks well: a function, an endpoint, a contained bug, a simple new project. With 32 GB or more the larger model completes almost every test project, but it is slow (about 10 tokens per second on an M1 Max). It does not replace the largest cloud models on the biggest jobs; in exchange your code never leaves the Mac. Ollocode requires an Apple Silicon Mac with at least 16 GB and can be downloaded from its page.

Frequently asked questions

Does it work without an internet connection?

Yes. Once the models are downloaded, both Ollama and Ollocode work offline. You only need a connection to download models and, if you allow it, to let the agent read a web page.

Does it work on Windows or Linux?

Ollama does, Ollocode doesn't: Ollocode is a macOS app and needs an Apple chip (M1 or later). Remote projects, however, can live on a Linux server reached over SSH.

Can I use a different model from the recommended one?

Yes, any Ollama or MLX model. The recommended ones are the most reliable among those measured for each amount of memory.