Asklemmy

44827 readers

901 users here now

A loosely moderated place to ask open-ended questions

Search asklemmy 🔍

If your post meets the following criteria, it's welcome here!

Open-ended question
Not offensive: at this point, we do not have the bandwidth to moderate overtly political discussions. Assume best intent and be excellent to each other.
Not regarding using or support for Lemmy: context, see the list of support communities and tools for finding communities below
Not ad nauseam inducing: please make sure it is a question that would be new to most members
An actual topic of discussion

Looking for support?

Looking for a community?

Lemmyverse: community search
sub.rehab: maps old subreddits to fediverse options, marks official as such
[email protected]: a community for finding communities

~Icon~ ~by~ ~@Double_[email protected]~

founded 5 years ago

MODERATORS

[email protected]

Can you self-host AI at parity with chatgpt? (lemmy.ml)

submitted 3 hours ago by [email protected] to c/[email protected]

4 comments fedilink hide all child comments

My office computer has a Ryzen 7 5700, RX 580x, and 32gb of ram. Running ollama with deepseekv2 or llama3 is much slower than chatgpt in the browser. Same with my newer, more powerful home computer.

What kind of hardware do you need to run with comparable responsiveness to chatgpt? How much does it cost? Presuming such hardware is commercial, where do you find it?

top 4 comments

sorted by: hot top controversial new old

[–] [email protected] 5 points 3 hours ago* (last edited 2 hours ago) (1 children)

Install LocalAI and ensure it’s using acceleration. It’s one of the best solutions we have at the moment.

Are you sure you’re not running these small models off of CPU and no acceleration? Because I’m running these small models pretty quickly. Nearly instant responses using a NVIDIA titanXP from a gaming rig I built in 2017 ish.

[–] [email protected] 4 points 3 hours ago

AMD is quite awful in this regard. Rn with my rx6650xt using Vulcan acceleration, I get the same speed as running on my r5 7600

[–] [email protected] 4 points 3 hours ago

You'd need basically a small server rack filled with datacenter GPUs. Expect mid to high 5 digit numbers.

But: running smaller models on a typical gaming GPU is quite doable.

[–] [email protected] -4 points 3 hours ago

What kind of hardware do you need to run with comparable responsiveness to chatgpt?

Generally you need between $8-10,000 worth of equipment to get relative responsiveness from a self-hosted LLM.