How I use a free local model as a delegated investigator for my AI assistant, trading wall-clock time for dollar cost — and what the research literature says about why this works.
Tag
Ollama
2 articles with this tag.
What I learned setting up local LLMs on consumer GPUs with Ollama — VRAM math, thinking-mode traps, and the cascade architecture I use for daily AI assistant work.