I have gone down this route but I'm afraid that 16K of context is nothing for agentic coding tasks. Even if you use a smarter model directly from an inference provider to plan all the work and then use your own 16GB GPU to execute the plan, managing those 16k tokens becomes unusable at some point even with a minimalistic harness like pi.
For other agentic tasks it's nice, and more than just the budget, it's the data sovereignty that you gain (imagine analysing tax forms, contracts, other legal documents, etc.) in my opinion.
Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech.
Big Tech models for anything else.
Thanks for writing and sharing. Since I also have a 4070 Ti Super, I'm always interested in hearing people's experiences with local models that fit 'middle-ground' hardware. I kind of feel left out because most posts I see seem to either be about clever ways of making older, lower-end cards usable, leveraging mega-CPU RAM (128GB) or how aweseome high-end GPUs like 4090/5090 can be.
A $1200 GPU you already own, typically, is how this is being seen. Maybe you already have a gaming PC, so this is 0 extra cost to you. All it takes is the power to run the GPU which would be minimal extra vs. the $20/mo cost.
Where I live electricity costs 30 to 40 cents per kWh, I think it's easy to go over 20 dollars per month. Also I find 20 dollars is really a small amount of money to spend on AI. AI fays itself back fast if it helps you with serious stuff.
A lot of quotas, reset timers, changed terms, and so on too through the subscription. I prefer to buy and use my tools, not adopt them as part of my lifestyle.
I bought my 4070 Ti Super in that magic ~month right in between the Crypto-Craze and the AI Bubble when good cards were pretty commonly available for around MSRP. I even found a small deal the day I was ready to buy, so paid $790 with 2-day delivery included.
Throughout the entire crypto-craze I'd been squeeking by gaming on an OC'd 1080 Ti, so when prices finally fell I was more than ready. Since I was getting interested in maybe playing around with AI soon, I spent up from my budget of around $550-$600 and am very glad I did since it's now clear I'll be squeeking by on this card for more years than I'd planned, just like the 1080 Ti.
Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech. Big Tech models for anything else.
"That stop being a complaint"? That's a strange construction.
I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.
Throughout the entire crypto-craze I'd been squeeking by gaming on an OC'd 1080 Ti, so when prices finally fell I was more than ready. Since I was getting interested in maybe playing around with AI soon, I spent up from my budget of around $550-$600 and am very glad I did since it's now clear I'll be squeeking by on this card for more years than I'd planned, just like the 1080 Ti.