Back

Getting 50 GB/S Back from the Apple Neural Engine

14 points3 dayseiln.github.io
Neywiny1 hour ago

Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?

eiln3 days ago

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.