Local AI

User avatar
antus
Site Admin
Posts: 10014
Joined: Sat Feb 28, 2009 10:34 am
cars: TX Gemini 2L Twincam 8psi
TX Gemini SR20 18psi
Datsun 1200 Ute
Subaru Blitzen '06 EZ30 4th gen, 3.0R Spec B
Subaru WRX 2007

Local AI

Post by antus »

I thought i'd start a thread about local AI, and new developments.

For now I just want to drop this. APEX and MTP are a massive improvement. All of a sudden getting very good results on local hardware with https://huggingface.co/mudler/Qwen3.6-3 ... X-MTP-GGUF

APEX is a quantization (compression) where different parts of the model get different amounts of compression. Good for higher quality at given size.
MTP is token prediction. Where the model predicts a number of tokens ahead in parallel and then verifies those tokens when it's ready to do the next one. Can be 2x or more speed increase with ~ 1Gb RAM requirement increase.
The above model in particular is so far running very well and fast on my local hardware.
Have you read the FAQ? For lots of information and links to significant threads see here: http://pcmhacking.net/forums/viewtopic.php?f=7&t=1396
8SecSleeper
Posts: 27
Joined: Sun Mar 12, 2017 5:57 am

Re: Local AI

Post by 8SecSleeper »

I started off using Claude, Gemini, then realized those models were useless 45 mins of usage when your locked out, then the weekly lockout.

Anyway I was gonna try local, then discovered the zen go plan, you pay $5 for first month, then $10 each month after. Gives you $30 usage of a range of open source models you could locally host.

I get sooo much done with mimo 2.5, if it gets stuck trying to solve something 2.5 pro instantly solves it, but it's very costly in comparison. But still way cheaper then the frontier models.

I estimate my cost to run locally is 6 cents per hour, But I'd need to build a seperate dedicated machine and put in basement so it doesn't add additional cooling requirements to my bill.

The math works out to 6 cents per hour locally vs about 7 cents per hour using Zen Go at the $10 cost which comes after the first month. Then add on the build cost. Also would need a way to keep the server at a very low cost power off and on when not in use, as the cost to leave it up 24/7 would make it far more expensive then using the zen go plan.
User avatar
antus
Site Admin
Posts: 10014
Joined: Sat Feb 28, 2009 10:34 am
cars: TX Gemini 2L Twincam 8psi
TX Gemini SR20 18psi
Datsun 1200 Ute
Subaru Blitzen '06 EZ30 4th gen, 3.0R Spec B
Subaru WRX 2007

Re: Local AI

Post by antus »

Wow, that's a big one! That would take something really special (read: really expensive) to run locally. 256Gb even at Q6 quant for the base version! That'd be like over $30k of hardware even from the second hand market. Could easily be triple that if you got quality. That pays for a lot of tokens in the cloud.
Have you read the FAQ? For lots of information and links to significant threads see here: http://pcmhacking.net/forums/viewtopic.php?f=7&t=1396
8SecSleeper
Posts: 27
Joined: Sun Mar 12, 2017 5:57 am

Re: Local AI

Post by 8SecSleeper »

antus wrote: Mon Jun 15, 2026 12:45 pm Wow, that's a big one! That would take something really special (read: really expensive) to run locally. 256Gb even at Q6 quant for the base version! That'd be like over $30k of hardware even from the second hand market. Could easily be triple that if you got quality. That pays for a lot of tokens in the cloud.
Yea in my research, it will continue to make sense to use api's while they are at a heavy discount, trying to prevent other models from gaining market share. Then when the pricing wars settle down, going to locally hosted would start to make alot more sense. And hopefully by that time, unified memory will be common, and you won't be stuck spending a fortune on old used GPUs.

Luckily these "open source" models are keeping the prices competitive, it doesn't make any sense at all to use frontier models for most tasks. Especially since I'm using my own custom ai harness and I don't use opencode, antigravity, cursor, claude code or even vscode. I've been slowly adding all those features that I would need into my own harness.

Attached is an example of my search ability, where it shows context to go with them, you can double click anywhere in the search results and the file is opened in the ide part to that part of the code.

You can fully edit, build, push, pull, sync.

Works with ssh connections, custom stuff like pterodactyl panel where the ai agent can fully create and deploy, debug custom plugins on game servers etc.

I could never afford to do this, using the google, openai, claude models. Well it wouldn't make a whole lot of sense at least.

I even have the timeline feature AI studio has, ability to roll back changes etc.

Automatic prompt reply, I can set it up to reverse engineer a ecu firmware, do 5 functions at time, auto audit when it's done and keep going, handoffs so the context stays reasonable.

Rosyln compiler in backend so anytime a change is sent to a c# file, it must past build checks, so Ai can't corrupt working files and go in fix it loops.
You do not have the required permissions to view the files attached to this post.
8SecSleeper
Posts: 27
Joined: Sun Mar 12, 2017 5:57 am

Re: Local AI

Post by 8SecSleeper »

What kind of results were you seeing in tk/s and what test hardware / harness?
Forgemaster98
Posts: 5
Joined: Thu May 14, 2026 12:29 am
cars: 1995 k3500 5.0 nv3500, 1998 c2500 5.7 4l60e, 2003 suburban 5.3 4l60e

Re: Local AI

Post by Forgemaster98 »

So I was lucky to have two beefy computers before the RAM-ageddon. I have an HP Z440 with an E5-2680 v4, 128 GB RAM, and an RTX 3060 12 GB that I use for small, fast models—normally Qwen3.5-9B-UD-Q6_K_XL with a 128k context at around 30 t/s for my everyday interactions as my chat model. And in my Dell T630 with dual E5-2697A v4 and 256 GB of RAM with 2x AMD Radeon Pro V620 32 GB for 64 GB of VRAM, I run the unsloth/Qwen3.5-122B-A10B-MTP-GGUF at Q4 128k context for coding and heavy reasoning at around 5-8 t/s, all through the Hermes agent, and it does very well for my needs. Thankfully, I do IT work and was into crypto for a while, so I had spare hardware lying around, although the T630 spends most of its time off due to power and heat. It can overwhelm the 12k BTU AC in my server room/shed."
hjtrbo
Posts: 353
Joined: Tue Jul 06, 2021 8:57 am
cars: VF2 R8 LSA
FG XR6T
HJ Ute w/RB25DET

Re: Local AI

Post by hjtrbo »

Really good intro primer video.
Setting up a local AI model can be intimidating, which is why I created this video. In this video I will show you everything you need to know about setting up local AI models from the basics of how local AI models work, to how to optimize local AI models for your hardware, and even how to use these local AI models in real world agentic environments.
https://www.youtube.com/watch?v=UngVdAsQEiU

And this is a good in-depth dive into AI principals & workflows for coding
https://www.youtube.com/watch?v=-QFHIoCo-Ko
User avatar
Ralmo94
Posts: 25
Joined: Thu Mar 12, 2026 10:00 pm
cars: 1995 towncar
1994 k1500
1998 k2500 7.4
2002 Tahoe 5.3
2004 Silverado 1500
2008 F250 5.4

Re: Local AI

Post by Ralmo94 »

One thing I have learned about using an AI for anything is what you get all depends on the prompt you give it, and how much context you add, and what context. They are trained to decide given this prompt in this context, what is the most predictable outcome the user is expecting. One thing that works realy well is asking an Ai how to prompt another Ai for what you are doing. They will craft a decent system prompt. I have used Chat GPT, used to be better, but now since they added adds in the chat, they nag you to upgrade constantly and demote you to a less powerful model in like two or three prompts, and the system is not made for coding, they want you to pay for that. Gemini is decent, but all the sudden you will notice context window flushed, and it is like the model lost almost all of the context in a thread and is just guessing what the last message was. Meta Ai is actually pretty decent, but as soon as you start talking about PCM bootloaders, or flash tools, security keys, or even bin dumps, you get a " I'm sorry I can't help with that, is there anything else I can help with" Claude is extremely good for a free tool, but if you have it doing real work, expect to do what it can in a 5 hour window and tell it to continue in 5 hours, but they are hard to beat. Recentley I started using DeepSeek, and I have noticed it is getting better, the Expert model can no longer search web or upload files, but the instant with think turned on works pretty good, if has a large context window, and I haven't ran out of usage once. Mistral AI now Vibe, is a lot better than it used to be, it now has a sandboxed file system like claude does, and the model seems a little smarter than it was, but sometimes it still seems not to be aware of all the tools it has. Usage also will also hit a cap, I think it resets in 2 hrs though? But it doesn't tell you, you just wont send your message to the model when you hit your limit. All of them are by far better than anything I could fit on any of my machines, I just have a few laptops and some business class desktops from over 10 years ago. Sorry for such a long first post here.
OEM boxes. Open questions. Closed source.
The engine knows what it wants. The ECU knows what it got. :hmm:
If the bootloader can read it, so can we. :gear:
The silicon doesn't lie — it's just not telling the whole story. :silent:
User avatar
antus
Site Admin
Posts: 10014
Joined: Sat Feb 28, 2009 10:34 am
cars: TX Gemini 2L Twincam 8psi
TX Gemini SR20 18psi
Datsun 1200 Ute
Subaru Blitzen '06 EZ30 4th gen, 3.0R Spec B
Subaru WRX 2007

Re: Local AI

Post by antus »

Yeah the context length is a major factor. For quite a while there with Claude I was just using the longest context I could, but then they changed their quota method so more context is more expensive. It used to get the hang of the code base and all of your rules but now if you keep going like that the usage stays quite high. Now you have to think about when to do a /compact and when to start a new thread and just drop the context. And meanwhile put more effort in to the memory system and give it clear short and shiny list of rules to remember so it can keep up with the right style and have the right background without a long conversation to get to that point. You can just start another conversation and pickup fresh but have it start with what it needs to be productive straight away.

Sadly localal is good for shorter coding tasks or regular usage, but the context possible on < $5k or < $10k of hardware just cant match the cloud where they do have that level of kit running full time with the costs shared across the entire userbase. Clearly AI will keep getting better and machines will keep getting bigger and one day it surely will be possible. Maybe 3-5 years? Almost certainly < 10 years.
Have you read the FAQ? For lots of information and links to significant threads see here: http://pcmhacking.net/forums/viewtopic.php?f=7&t=1396
User avatar
Ralmo94
Posts: 25
Joined: Thu Mar 12, 2026 10:00 pm
cars: 1995 towncar
1994 k1500
1998 k2500 7.4
2002 Tahoe 5.3
2004 Silverado 1500
2008 F250 5.4

Re: Local AI

Post by Ralmo94 »

One thing with claude, make sure you use the skills feature for stuff you use it for all the time. You can just tell it you want it to create a skill. Basically instead of user preferences and thread context, and guessing what to do to satisfy the request, it looks at the skill file as instructions and follows them pretty well. I tried using that feature with Grok, and it kept doing stuff the skill told it not to, I finally asked it why bother setting a skill up if you aren't going to follow it? It obeyed and apologized, then a little while longer did it again, then i was out of tokens for the Grok ridiculous 12 hour window lol. Never had that problem with Claude, they train there models differently than everyone else. If you haven't tried Codex yet, you should, if you get the desktop app you can use it free with a chat gpt account. Token allotment used to be per week, but last time I used it they changed to monthly. It can run terminal commands on your machine, write and patch files, and test them, and iterate on them until it works.
OEM boxes. Open questions. Closed source.
The engine knows what it wants. The ECU knows what it got. :hmm:
If the bootloader can read it, so can we. :gear:
The silicon doesn't lie — it's just not telling the whole story. :silent: