

Context limit is not really a problem on local models. Qwen3.6 can do up to ~256k tokens, it’s not that far from what things like cursor uses. grok in cursor have exactly 256k tokens for example. Also, you can use opencode with custom config, where you need to set trimming close to that number.
That being said, quality wise it still kinda shit and looses to paid models. I think you need something like GLM 5.2 to compete, which needs 228gb at lowest 1bit quant (not including context size, which can be up to 1m tokens, so you can round up requirements for RAM up to 256gb), so yeah, the only thing that limits you locally — the fact that “AI” companies bought all supply of focking ram and we can’t afford any.









I have no clue why that triggered people so much. Never stated that I have any problem with that, nor that I need any help solving it. I leave my apartment like, once per month and cause living in EU - there no need to “travel” further than opposite street to buy anything I may need.
On top of that, point was that it’s issue for other people in general. Explain how to setup tailscale to your grandma, I dare you.