62
you are viewing a single comment's thread
view the rest of the comments
[-] eager_eagle@lemmy.world 7 points 1 month ago* (last edited 1 month ago)

yeah, that is pretty insane at 17k tokens per second. It feels as if answers are already cached just waiting for you to prompt.

https://chatjimmy.ai/

If they manage to fit larger and more recent models in it, it could greatly improve the energy requirements to run these models.

The recent diffusion LLM from google is also really exciting, and I think they might become the new architecture of choice for these models when running on general purpose GPUs - especially consumer cards, which are usually memory constrained, but computing-capable.

this post was submitted on 14 Jun 2026
62 points (100.0% liked)

Technology

86748 readers
2411 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS