2

Meta plans to start production of its in-house AI chip, called Iris, in September as part of a major expansion of its AI infrastructure. The company aims to double its computing capacity to 14 gigawatts by 2027, reduce reliance on external chip suppliers, and lower AI operating costs.

1

The cheaper AI tokens get, the more expensive memory becomes. Inference costs fell 280-fold in two years; enterprise AI spending tripled anyway, and DRAM contract prices just posted a record 90-95% quarterly jump. Jevons saw this pattern in 1865 with coal. A look at what actually rebalances the chip market — and why the answer is fabs in late 2027, not smarter models or ASICs.

2
3

Spending $2,700 on 8TB storage for a $700 game console?

SanDisk announced the officially licensed Optimus GX PRO 850P M.2 NVMe SSD for PS5 and PS5 Pro, offering up to 8TB of PCIe 4.0 storage with a custom heatsink and high performance, but its extremely high pricing has quickly become the main talking point...

#SanDisk #SSD #PlayStation #Gaming

1

Amazon is no longer just building AI chips for internal AWS use. It is preparing to sell its Trainium AI chips directly to external customers, turning itself from a cloud provider into a direct AI hardware competitor to NVIDIA. The bigger story is that hyperscalers (Amazon, Google, Microsoft, Meta) increasingly want to control their own AI infrastructure instead of depending entirely on NVIDIA.

3

AMD bought MEXT, a startup whose software makes cheap NAND flash act like DRAM, so a server runs on far less of the expensive stuff. With DRAM contract prices up ~90%+ this year, the logic isn't subtle. How predictive memory tiering actually works?

1

AMD Ryzen AI Max+ 395 vs Nvidia DGX Spark vs Apple Mac — plus when a GPU tower still beats all three. A practical hardware guide for IT managers, developers, and small-business owners weighing a local LLM machine.

The short version: there’s no single winner. The right pick comes down to three numbers — how big a model you need to run (memory capacity), how fast it has to run (memory bandwidth), and which software you depend on (CUDA, ROCm, or Metal). A discrete-GPU tower is fastest but hits a VRAM wall; AMD’s Strix Halo mini PCs give the most memory per dollar on Windows; Nvidia’s DGX Spark adds the CUDA stack at a premium; Apple’s Macs offer high bandwidth and silence without CUDA. An appendix at the end collects what early buyers of the AMD “lunchbox” are actually reporting.

How to Choose the Best Mini PC for Local AI

7

Interesting...The AI race is moving so fast that they can not wait for a permanent building

1

Arm challenged it on two fronts at once. At Computex 2026, NVIDIA's RTX Spark put a 20-core Arm CPU + Blackwell GPU inside mainstream Windows laptops, while Arm's first data-center chip — the AGI CPU — hit mass production with Meta, OpenAI and Oracle aboard. So what happens to x86 servers and parts out there now? Here is the full article: Arm vs x86

3

Dell just reported $16.1B in AI server revenue — up 757% year-over-year — with $24.4B in new orders and a $60B full-year forecast. The bigger story: enterprise AI has crossed from pilot projects into production infrastructure. The new constraints aren't models or software. They're power, cooling, memory, and rack density.

4

"Intel Arc G-Series represents years of focused innovation and a deep commitment to gaming. It delivers uncompromising PC performance in the palm of your hand, combined with the console-like accessibility and immediacy gamers expect. With cutting-edge graphics technologies like XeSS 3 and breakthrough efficiency for longer unplugged play, Intel Arc G-Series proves that while others make tradeoffs, gamers don't have to."

2
submitted 5 months ago by alexbsr@lemmy.sdf.org to c/AINews@lemmy.world

The key takeaway isn’t just compression—it’s where the bottleneck shifts. KV cache has been dominating memory footprint in long-context inference, so reducing it changes the cost structure significantly. But it doesn’t remove the constraint entirely:

You’re trading memory bandwidth for additional compute (de/quantization isn’t free) Model weights and activation flows still sit in high-bandwidth memory At scale, efficiency gains often trigger more usage (classic Jevons paradox)

One implication that doesn’t get discussed enough: this could extend the useful life of existing GPUs (A100/H100 class) for inference workloads, especially for long-context applications.

Curious how people here see this playing out in production systems—does KV cache compression meaningfully change your infra decisions, or just shift optimization elsewhere?

Will Google’s TurboQuant AI Compression Finally Demolish the AI Memory Wall?

[-] alexbsr@lemmy.sdf.org 4 points 10 months ago

I see, thanks! I will be more cautious. But what I am selling is physical computer memory, not a website, not scamming.

[-] alexbsr@lemmy.sdf.org 3 points 10 months ago

Yes, if such post is not allowed, I can delete it. thanks.

[-] alexbsr@lemmy.sdf.org 2 points 10 months ago

what is it?

view more: next ›

alexbsr

joined 2 years ago
MODERATOR OF