Local AI

User avatar
antus
Site Admin
Posts: 10012
Joined: Sat Feb 28, 2009 10:34 am
cars: TX Gemini 2L Twincam 8psi
TX Gemini SR20 18psi
Datsun 1200 Ute
Subaru Blitzen '06 EZ30 4th gen, 3.0R Spec B
Subaru WRX 2007

Re: Local AI

Post by antus »

I started out with codex to see how it was while everyone else was on the claude bandwagon. At that time I thought it was pretty good, but eventually swapped over to claude and so far have not gone back. Well, my day job does use openai, so I have codex through that, but it just frates the hell out of me these days, going around in circles and repeatedly telling me something that I know that isn't relevant. I have to get real shirty with it for it to get the point that I understand and to stop nagging. I am not going to pay money for a machine that is supposed to be a tool to keep repeating itself to impress point that we have well moved past. I use claude in vscode with the plugin and that is great with tools and the shell, and on the topic of this thread, so far ive found the open version of google gemini - gemma 4 is the most solid for tool use and MCP locally. But without enough ram for huge context, it can still only handle jobs up to a certain level of complexity. Basic to medium software development or RE yes, but you will hit the limits even on a 31b model.
Have you read the FAQ? For lots of information and links to significant threads see here: http://pcmhacking.net/forums/viewtopic.php?f=7&t=1396
User avatar
Ralmo94
Posts: 25
Joined: Thu Mar 12, 2026 10:00 pm
cars: 1995 towncar
1994 k1500
1998 k2500 7.4
2002 Tahoe 5.3
2004 Silverado 1500
2008 F250 5.4

Re: Local AI

Post by Ralmo94 »

You might copy some of that when it happens and ask another ai to craft a system prompt to avoid that, save it as a txt and start every new thread by pasting that in. Codex doesn't really talk much when I use it, I mostly have it write patch and debug stuff, but it seems like it has to figure some things out it should already know. Chat GPT - I hear ya with rant type stuff, one I got from it frequently was, you accidentally did that the right way, you just didn't know it. Who said I didn't know it and it was accidentally? lol I think a lot of their training material comes from open forums, so who knows what all the model has read lol
OEM boxes. Open questions. Closed source.
The engine knows what it wants. The ECU knows what it got. :hmm:
If the bootloader can read it, so can we. :gear:
The silicon doesn't lie — it's just not telling the whole story. :silent:
8SecSleeper
Posts: 27
Joined: Sun Mar 12, 2017 5:57 am

Re: Local AI

Post by 8SecSleeper »

antus wrote: Sat Jun 20, 2026 7:45 am I started out with codex to see how it was while everyone else was on the claude bandwagon. At that time I thought it was pretty good, but eventually swapped over to claude and so far have not gone back. Well, my day job does use openai, so I have codex through that, but it just frates the hell out of me these days, going around in circles and repeatedly telling me something that I know that isn't relevant. I have to get real shirty with it for it to get the point that I understand and to stop nagging. I am not going to pay money for a machine that is supposed to be a tool to keep repeating itself to impress point that we have well moved past. I use claude in vscode with the plugin and that is great with tools and the shell, and on the topic of this thread, so far ive found the open version of google gemini - gemma 4 is the most solid for tool use and MCP locally. But without enough ram for huge context, it can still only handle jobs up to a certain level of complexity. Basic to medium software development or RE yes, but you will hit the limits even on a 31b model.
I played with gemma a bit, I think it's not worth using at this point, It's nice it's available but until we have system wide high speed ram, it will not be worth trying to mess with. Because even if you have enough space to host the model, once context starts leaking into ram space, the speed comes down so much it's not worth running.

However, smaller stuff like embedding models, voice to text stuff like that is really easy to run and get high speed locally.

They are working on, have working unified ram for windows, mac already has it on newer systems. Once that's standard, gpu prices won't make sense anymore, at the same time ram prices will be terrible. Which they already are. My 64gb ddr5 was $500 new, dropped to around $200, now it's $1050-$1200. With unified ram, this will easily get worse till there is more supply and the bubble burst, kinda like when pc's were less common.
User avatar
AngelMarc
Posts: 612
Joined: Sat Apr 08, 2023 11:23 am
cars: A CB450 running to 8,000RPM with a P59.

Re: Local AI

Post by AngelMarc »

Ralmo94 wrote: Sat Jun 20, 2026 5:42 am I have used Chat GPT, used to be better, but now since they added adds in the chat, they nag you to upgrade constantly and demote you to a less powerful model in like two or three prompts
I had better luck with the "less powerful" models sometimes.
I was going for rock dumb, fast code though.
Don't stress specific units.
User avatar
Ralmo94
Posts: 25
Joined: Thu Mar 12, 2026 10:00 pm
cars: 1995 towncar
1994 k1500
1998 k2500 7.4
2002 Tahoe 5.3
2004 Silverado 1500
2008 F250 5.4

Re: Local AI

Post by Ralmo94 »

AngelMarc wrote: Sat Jul 04, 2026 12:16 pm
Ralmo94 wrote: Sat Jun 20, 2026 5:42 am I have used Chat GPT, used to be better, but now since they added adds in the chat, they nag you to upgrade constantly and demote you to a less powerful model in like two or three prompts
I had better luck with the "less powerful" models sometimes.
I was going for rock dumb, fast code though.
Yeah instructions following is worth more than an intelligent model I have learned. If you are working with firmware it is good to be able to anchor the ai to a skill with clear rules about embedded firmware and chip specifics from the manufacturer manuals. The two best free options I have found for this are Claude and mistrial, they both will follow the rules well. For fast code I usually use deepseek
OEM boxes. Open questions. Closed source.
The engine knows what it wants. The ECU knows what it got. :hmm:
If the bootloader can read it, so can we. :gear:
The silicon doesn't lie — it's just not telling the whole story. :silent: