• ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
    link
    fedilink
    arrow-up
    0
    ·
    11 hours ago

    I expect we’ll start seeing stuff like Taalas where they print the model to the chip and other specialized chips like Xuantie C950 going forward. Neither of these requires DRAM, and Taalas is particularly clever since they just print the model right to an ASIC chip. So, the whole renting out LLMs business model isn’t going to last long I suspect.

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      0
      ·
      9 hours ago

      So a repeat of the crypto crash for graphics cards when ASICs ate their lunch. Mind you, that’s only for inference (although a super fast QWEN 3.8 would meet a lot of peoples needs).

      The argument for datacentres is for training the models, but then they’ll need to prove that they haven’t hit a diminishing returns wall, which will be hard if, as seems likely, they have. Also the Chinese have been doing it in a cave, with a box of scraps (figuratively), and gotten at least 90+% as good results.

      Seems like the recent advances have been in the frameworks, which don’t need no stinking (literally if fossil fueled) datacentres.