228
Apple M7 Ultra Chip Planned With Up to 1.5 TB of Unified Memory
(www.techpowerup.com)
This is a most excellent place for technology news and articles.
I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence.
That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE)