Another reason to self host your own AI

SuspiciousCarrot78@aussie.zone · edit-2 30 days ago

Another reason to self host your own AI

Alavi@programming.dev · edit-2 28 days ago

I started working toward self hosting LLM for my small company using ollama and opencode as agent But I realized a good model like GLM 5 requures 250GB of RAM and 24GB vram with a 4080?? I dont know, this is what the LLM told me itself.

I ended up using qwen-code2.7-7b-16k.

Currently the best thing I have is my laptop, 16GB ram, i7 9750H gtx1650

How do you guys selfhost? What models do you use that are actually good?

SuspiciousCarrot78@aussie.zone · edit-2 27 days ago

I mean…that entirely depends on your use case - and I hate saying that. For me and what I do, Qwen SLM (esp Qwen3-4B 2507 instruct and Qwen3.5-2B) are exceptional. But I’m not trying to do Claude at home.

Best bet? Spend $10 on OpenRouter and try different models. In a head to head with ChatGPT 5.4 mini (excellent for coding BTW), I’ve found Qwen 3.5 27B more than able to hold its own for coding tasks…IF you narrowly gate it/confine it. The last batch of Qwen’s really are something. Dunno about the 3.7 series.

Having said ALL that, I’m really tempted to go back in time and code myself a deterministic expert system, with user updatable knowledge cascade, tool calling and a minimal amount of Markov chain word garnish for flavour. I think we use to just call that “a program” lol.

Really tempted actually, because if 50% of llm use case is basically Super Google but not shit…well, I can make that myself. I just need to point my autism at it.

PS: this might help

https://www.youtube.com/watch?v=0AqpaFm11oI

Alavi@programming.dev · 27 days ago

Qwen 3.5 24B is way too large for my specs. I’m barely running qwen2.5 7B

SuspiciousCarrot78@aussie.zone · edit-2 27 days ago

Hmm…it runs on a 1060…it’s a MoE not a dense. 24B is even lighter. Worth a shot.

https://www.youtube.com/watch?v=8F_5pdcD3HY

Else, if youre looking for a coding model (??) something like Sara or fara might suit

https://huggingface.co/microsoft/Fara-7B

brucethemoose@lemmy.world · 30 days ago

Yeah.

It’s not even about efficiency, really, but independence from corporations, privacy, and principle. Kind of like Lemmy.

Auli@lemmy.ca · 29 days ago

Sure but all these self hosted ais are still done by companies who used massive amounts of power and water to train it.

KatherinaReichelt@feddit.org · 29 days ago

Which is an interesting dilemma: Those AIs are already trained. That power and water was used. If you use them, you will not pollute anything. But you may encourage those companies to train another AI

brucethemoose@lemmy.world · 29 days ago

No.

Even the biggest open weights models are trained on pennies compared to OpenAI and Claude. They just don’t have the hardware to be so wasteful.

In fact, the Nvidia GPU ban was the best thing to ever happen to “small” AI devs. It made them thrifty.

Noxy@pawb.social · 29 days ago

not gonna self host bullshit that wastes resources and makes me dumber.

toor@lemmy.world · 29 days ago

Me, looking at my Jellyfin server…

Oh. Ok.

Hiro8811@lemmy.world · 30 days ago

You’re still paying for electricity and a big part of the world is in a electricity crisis. “AI” has few real uses and LLMs are not one of them.