My office computer has a Ryzen 7 5700, RX 580x, and 32gb of ram. Running ollama with deepseekv2 or llama3 is much slower than chatgpt in the browser. Same with my newer, more powerful home computer.

What kind of hardware do you need to run with comparable responsiveness to chatgpt? How much does it cost? Presuming such hardware is commercial, where do you find it?

  • JoYo
    link
    fedilink
    English
    8
    edit-2
    3 months ago

    It’s all dependent on VRAM. If you can load the distilled models with your GPU without maxing out your VRAM it will run just as fast as any server farm.

    RX 580x

    It looks like your video card only has 8 GB of VRAM. That will be your bottleneck.

      • JoYo
        link
        fedilink
        English
        23 months ago

        yah that’ll do it too. ive got a 6800xt which isn’t technically supported but it works well.