Back

Platform-Independent SIMD in Go

283 points7 hoursgo.dev
ImJasonH • 4 hours ago

https://imjasonh.github.io/playground/palette-swap/ swaps colors in a provided image in wasm, entirely locally in your browser, to benchmark portable SIMD vs non-portable archsimd vs non-SIMD.

Portable SIMD is ~11% slower than non-portable SIMD in this case, but both are ~5x faster than non-SIMD.

u8 • 5 hours ago

This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.

physicsguy • 3 hours ago

People were definitely asking for it.

typical182 • 2 hours ago

It's been discussed for a long time, and the related proposals were heavily upvoted, including various older proposals.

As I understand it, part of the reason it took a while is that the core Go team was generally of the opinion that doing user-facing SIMD APIs the right way was to design a high-level, cross-platform API that would stand the test of time, and that was then punted a few times given its complexity and need to do other things.

Part of what helped the current approach take off was switching to a philosophy of designing a lower-level architecture-dependent API first (the 'simd/archsimd' package), and then later doing a higher-level portable API (the 'simd' package, which is topic of this blog post).

That two-level approach I think also gave some additional freedom for the design and implementation of the friendlier / high-level 'simd' package, including because the lower-level 'simd/archsimd' package is available for people who need or want to drop down.

It's a nice design.

senderista • 2 hours ago

Layered API design is a great way to resolve ergonomics/performance tradeoffs.

nonethewiser • 1 hour ago

This is kind of the opposite of Go. Not giving people what they are asking for.

There are pros and cons of course. You don't have 17 different ways to iterate over an array, so that's nice. But you also went 13 years without generics, despite them being one of the most requested features, because the designers didn't want that complexity inside Go.

Overall I think Go is better for this philosophy but there are times where the language is clearly written more for its maintainers than it's users.

andrewstuart2 • 50 minutes ago

Some of the concerns around generics and why it took so long were for the users as well. One of the biggest draws to Go has always been that you get the performance of a compiled language and yet compile times are so low that it can feel like you're developing with an interpreted language. The design of generics needed to maintain the compile time advantage or else it wouldn't feel like Go any more.

pjmlp • 3 hours ago

Mostly safe, contrary to other safer languages, Go memory model doesn't prevent data tearing.

__s • 5 hours ago

go data races aren't memory safe

tptacek • 3 hours ago

That's not what "memory safe" means. "Memory safe" is a term of art meaning "not susceptible to memory corruption exploits", like stack and heap overflows, UAFs, and type confusion. Last I checked, there are essentially no non-contrived memory corruption exploits for Go programs; the best you get are people demonstrating register control on contrived programs.

The definition I'm giving is the same as the ISRG's definition at MemorySafety.org. It's the thing everybody is talking about when they talk about memory safety.

The claim being made here is "big if true", because it would imply a lot more languages than Go "aren't memory safe", despite decades without memory corruption exploits.

monocasa • 22 minutes ago

The exploits aren't the only issue; they just get a lot of air time.

It's remarkably easy to segfault Go applications with data races. Any object with multiple words (so a slice that's an array pointer and a size, or a fat pointer with the object pointer and the vtable pointer), can be read in an inconsistent state from two threads which can cause out of bounds reads and writes. And this comes up all the time with how heavily the language encourages concurrency.

It's just difficult to actually exploit because of other considerations that practically add a lot of runtime entropy.

0c3ca83 • 3 hours ago

Yeah, we get it, you performatively hate go.

ngrilly • 3 hours ago

Yes, but in practice they are extremely hard to exploit. It has been discussed extensively here on HN and in other forums.

shikck200 • 4 hours ago

That does not make sense to me. Go is memory-safe, but it does not guarantee data-race freedom.

So whats your point here? Haskell?

SupLockDef • 3 hours ago

Probably Rust, that's always Rust with this kind of comments...

+1
asdf88990 • 3 hours ago
beltsazar • 3 hours ago

> That does not make sense to me.

You said that because you assumed Go is memory-safe in all conditions.

> Go is memory-safe

Yes, but only if there's no data race.

Go is not like Java. Java doesn't guarantee no data race, but when it happens, it's still memory-safe.

+2
shikck200 • 2 hours ago
simonask • 3 hours ago

There is no memory safety without freedom from data races. One is a prerequisite of the other. This is why languages like C# throw exceptions on unsynchronized concurrent access to some container types, and treat all property accesses as atomic.

+1
typical182 • 3 hours ago
+1
shikck200 • 3 hours ago
amelius • 3 hours ago

I suppose that if you create a map in one thread, and then access it from another thread, then that might cause segmentation faults? Because map is a type that is implemented in C.

OutOfHere • 5 hours ago

(removed)

typical182 • 4 hours ago

Go is broadly considered to be a memory safe language.

See for example comments from tptacek like:

https://news.ycombinator.com/item?id=43335748

https://news.ycombinator.com/item?id=46028232

https://news.ycombinator.com/item?id=44672371

(The gist: memory safety is a term of art coined by security practitioners. Go, Python, Rust, Java, others: memory safe. C/C++: memory unsafe. Periodically, people in different slices of industry or academia come up with new definitions of memory safety that declare Rust or Go or other languages to be memory unsafe, but that is not by the broadly accepted definition across industry.)

mitxela • 4 hours ago

Rust does allow you to overflow buffers, confuse types, and duplicate mutable pointers in safe code. See cve-rs.

+3
simonask • 3 hours ago
saagarjha • 3 hours ago

He’s very wrong about this. Just because ‘tptacek posts a lot and did security once upon a time does not make him “broad consideration”.

shikck200 • 4 hours ago

You can write unsafe code in Go (import unsafe), but then, you can do the same in Rust. Unsafe code is not the default, and in day to day Go i rarely see the use of the unsafe package.

iambvk • 3 hours ago

What he probably means is data-races in go can result in memory/type unsafe accesses -- I suspect, likely due to slice types -- not sure if that is true/false.

+2
shikck200 • 3 hours ago
seki285 • 4 hours ago

No idea why you're getting downvoted for true statement. Without a ? like in C# you're always at risk of a nil pointer being dereferenced

bel8 • 4 hours ago

> like in C# you're always at risk of a nil pointer being dereferenced

That throws a NullReferenceException

+2
seki285 • 4 hours ago
mshockwave • 2 hours ago

Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision

janwas • 1 hour ago

We pioneered this in Highway and shared some advice on the API. Great to see this decision taken :D

melodyogonna • 1 hour ago

How so? I imagine you'd still want to constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.

Mojo has an even more portable simd[1] type that isn't just generic over length but also over type. In my opinion it is almost always better to specialize for each platform and use portable implementation as fallback. It's a shame that just very few languages support Zig-like comptime, because it would be excellent for specializations without introducing runtime penalties.

1. https://mojolang.org/docs/std/simd/SIMD/

sixdimensional • 2 hours ago

I did some testing with the experimental SIMD on a project I was doing to make speech-to-text and text-to-speech models run natively in Go (with CGO_ENABLED=0, so no C depenencies), and testing non-SIMD w/ SIMD.

I don't have formal benchmarks for that, but I can anecdotally say the SIMD work made a measurable improvement in the performance of the calculations vs. just plain Go. I'm very optimistic about how these improvements will help make the Go runtime an even better target for more of these types of work going forward, especially since it is cross-platform.

beached_whale • 4 hours ago

C++ is getting std::simd in the latest version and I am all aboard writing the vectorization with the least amount of intrinsic builtins I am able to. Even if not optimal, it's far better than the scalar ops.

reactordev • 3 hours ago

Seconded!! This doesn’t really help the well established codebases much that are already doing this on a platform specific path but in general this is much appreciated for the future.

beached_whale • 2 hours ago

Write it once with N errors, not N*M errors :)

melodyogonna • 50 minutes ago

Very neat, and comes pretty close to how Mojo handles portable SIMD.

It's great to see two of my favorite languages finally making SIMD easy to use. It's such low-hanging fruit for performance, yet somehow languages have ignored it for years. Portable SIMD, even with some performance penalty, still beats scalar computation whenever vector operations are needed. Yet language implementations always seemed to assume that hardware-specific SIMD APIs were the only way to go. That did nothing but make SIMD unusable excepting special cases where performance is absolutely critical, rather than just something anyone can use in day to day programming.

qprofyeh • 7 hours ago

This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.

pjmlp • 6 hours ago

Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.

So yeah, also appreciate having Go in the group instead of manually having to write Assembly.

However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.

stingraycharles • 6 hours ago

Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.

I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.

pjmlp • 5 hours ago

Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.

However even then, it depends on how the libraries API surface looks like.

vlod • 3 hours ago

You probably weren't looking for a tutorial about SIMD, but just in case you were interested, Mitchell [0] did one recently that got on HN [1]

[0] Mitchell Hashimoto: "Everyone Should Know SIMD" https://mitchellh.com/writing/everyone-should-know-simd

[1] https://news.ycombinator.com/item?id=49010648

setr • 18 minutes ago

I definitely was so... thanks!

janwas • 1 hour ago
Thaxll • 6 hours ago

With AI I'm pretty sure SIMD will be easier to integrate when necessary.

stingraycharles • 5 hours ago

But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.

Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.

preisschild • 4 hours ago

In 2026 if you are not doing A with AI you are doing it wrong /s

pjmlp • 6 hours ago

With AI, I expect it to eventually be good enough for us to finally have 5 GLs, so it won't really matter.

"CGO 2022 Keynote: Compiler 2.0"

https://www.youtube.com/watch?v=w_sX9aZoZxg

abirch • 6 hours ago

Vectorizing computations has been Matlabs secret sauce.

KeplerBoy • 6 hours ago

Does matlab these days do stuff like JIT operator fusing to avoid memory roundtrips and take advantage of FMAs?

abirch • 5 hours ago

Yes: https://www.mathworks.com/help/fixedpoint/ref/half.fma.html

I'm grateful that Go a non-proprietary language offers these features.

mastermage • 6 hours ago

Julia does that too.

metaltyphoon • 26 minutes ago

For God sake, add a syntax highlighting on the official page! Otherwise this is awesome

rcarmo • 52 minutes ago

I am using Go assembly for SIMD very heavily in https://github.com/rcarmo/go-pherence, this is just icing on the cake.

vira28 • 4 hours ago

This will welcome more database/warehouses to be written in Go.

Personally I will implement it in https://github.com/viggy28/streambed

pjmlp • 3 hours ago

They could have used third party packages or Assembly directly.

This naturally is an easier way.

vira28 • 32 minutes ago

You're right. I could have but this encourages me to seriously consider it.

ghusbands • 2 hours ago

> The new simd package hides these differences by removing fixed-size vectors from the type system, and by only supporting those operations that are in the intersection of all the different platforms, and fills gaps in the intersection with efficient emulation in terms of other SIMD instructions.

The intersection would be the operations supported by all platforms and so would not have gaps.

vlovich123 • 5 hours ago

> The interface conversion and type switch look like they should be inefficient, but the compiler-side implementation of simd specializes code and optimizes away the type switch.

I don’t understand this - how is it able to if the same go binary might run on unknown types? I’m assuming what it means is that the switch is implemented efficiently due to CPU branch prediction? I know fearless SIMD is doing cool stuff with static dispatch so that the feature set is checked just once at program start - is that what it means it’s doing under the hood? Very unclear.

Scaevolus • 5 hours ago

It creates multiple versions of functions referencing SIMD and lifts the dispatch switching cost to their callers.

> The AST rewrite creates multiple specialized copies of functions, variables, and types that mention simd types, where simd types are replaced with references to size-specialized types in simd/internal/bridge. Each of these bridge types is defined as an archsimd type, but with a restricted set of methods. The specialized functions, variables, and types acquire a suffix of the form @simdNNN, where NNN is either a vector length (128, 256, or 512) or 0, indicating emulation. Functions that mention simd internally, but not in their signature, are converted to wrappers that switch on the SIMD level detected at program start, and call the appropriate specialized version of that function. Specialized functions call other specialized functions directly without dispatch overhead (and perhaps with inlining). This rewrite strategy was chosen as a compromise between code duplication and SIMD performance; the overhead is hoisted as high as necessary to avoid dispatch within SIMD computations, but not higher. If SIMD dispatch appears “too low” in a computation, a gratuitous mention of a simd type will move it upwards, as in this example:

keel_dev • 3 hours ago

[flagged]

physicsguy • 6 hours ago

Oh this is great, it was one of my biggest bugbears about Go since you almost always have to link C/C++ code to get the appropriate performance.

The one negative I'd say is that often autovectorisation is 'good enough' and this doesn't really tackle that gap.

typical182 • 6 hours ago

FWIW, there is some pretty substantial autovectorization work that is already in-flight for the Go compiler.

There's a CL stack here:

https://go.dev/cl/791740

It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.

It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)

tgv • 6 hours ago

As a first step, it might be possible to write a linter rule that rewrites suitable numeric loops to SIMD. There are already rules to rewrite several loop types, so that should be doable.

pjmlp • 6 hours ago

The poor Assembler and the unsafe package forgotten in the corner.

While reaching out to CGO is the easier way, it doesn't mean it is the only tool available in Go.

fatty_patty89 • 6 hours ago

The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough

pjmlp • 3 hours ago

Use Assembly instead of CGO, isn't that scary, back in the 8 bit days we were coding Assembly aged 10, on our Spectrum, C64, Atari, Apple, Acorn, MSX,....

victorbjorklund • 5 hours ago

Why is that the case? I don’t know low level programming so why is Go limited in interop with C?

fatty_patty89 • 3 hours ago

its not limited but it has overhead because of the memory model of go doesnt match the C one so there has to be some sort of rerodering being done, that's what i understood atleast, and theres also the go concurrency

Am4TIfIsER0ppos • 35 minutes ago

How many "functions" compile to movd?

sharktheone • 3 hours ago

I hope portable simd will be stabilized some time in rust :/

karolist • 6 hours ago

Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.

shevy-java • 3 hours ago

Rust kind of seems to have overtaken Go in momentum recently. I wonder if Go will do well in, say, two years from now on.

jessechili • 40 minutes ago

[flagged]

neonsunset • 3 hours ago

[dead]

chrisjj • 5 hours ago

> Go 1.26 and 1.27 include experimental APIs for Single Instruction Multiple Data (SIMD) operations.

You'd think these people would know the meaning of API, no?

cloudfudge • 4 hours ago

One wonders what overly-narrow definition of API you're stuck on.

chrisjj • 3 hours ago

One wonders why you wonder.

https://en.wikipedia.org/wiki/API

tredre3 • 1 hour ago

I suspect you stopped reading at web services on your link, but API is indeed the correct word to describe a set of functions from a library (built-in or not). If you disagree perhaps you should share your preferred term here?

> The term API is often used to refer to web APIs, which allow communication between computers that are joined by the internet. There are also APIs for programming languages, software libraries, computer operating systems, and computer hardware.

chrisjj • 14 minutes ago

[delayed]