Back

U.S. Department of Energy Launches the Genesis Open Models Initiative

92 points4 hoursgenesisopenmodels.anl.gov
firasd2 hours ago

Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?

Ah but Mira Murati's new Inkling is Apache 2.0

But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC

ipsum22 hours ago

There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.

embedding-shape1 hour ago

Laguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).

walrus0135 minutes ago

There was an obvious problem in the original release, they re issued it after like a week with the reasoning looping supposedly fixed.

behnamoh59 minutes ago

No it doesn't follow instructions and is substantially slower than ds4.

kadoban26 minutes ago

It's a lot smaller, and runs (quantized) on a 3090 quite well. Ds4 flash 0731 you're talking about? It's great but it's much harder to run locally.

+1
jauntywundrkind52 minutes ago
kadoban1 hour ago

Yeah I think it got bad press because the chat templates (or something?) were messed up on first release, but I've been using a quant of it and it's a powerhouse, better than qwen 3.6 27b for local on a 3090, which is saying a lot.

firasd2 hours ago

Just looked into some Nemotron stats

Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27

On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia

coder5431 hour ago

I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.

At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.

The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.

Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.

Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.

(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)

written-beyond2 hours ago

Don't forget IBM

walrus0136 minutes ago

Laguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible.

https://huggingface.co/unsloth/Laguna-S-2.1-GGUF

loeg2 hours ago

I would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).

solomatov2 hours ago

Which remarks? Could you share a link?

wmf2 hours ago

Also Nemotron and Arcee.

mistrial91 hour ago

review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.

connorbrinton1 hour ago

Laguna S 2.1 is another fairly impressive-for-the-size American open model

andsoitis1 hour ago

Does Europe have an equivalent program?

shakna22 minutes ago

[delayed]

behnamoh58 minutes ago

[flagged]

plazmatic41 minutes ago

[dead]

an0malous1 hour ago

Do all these models have any significant architectural differences or training data sources? What are the factors going into the diversity of their performance?

ux26647856 minutes ago

The article posted is basically entirely about that.

Smith422 hours ago

What would the selected participants get from this? Looks like there is no offer of funding?

datlife1 hour ago

This is refreshing considering all the FUD (mostly from 1 frontier lab) happening around Open weight models.

Thegn2 hours ago

“Gomi” is the Japanese word for garbage. Gotta wonder if someone has a sense of humor…

greggsy1 hour ago

The Australian Liberal Party (basically our version of conservative republicans) proposed the National Energy Guarantee policy in 2017, which inevitably failed due to the media and public’s relative literacy and tendency to turn policy names into acronyms.

andsoitis1 hour ago

I wonder why it took so long.

dmix1 hour ago

Mostly because it's generally a bad idea for government to try to compete with a brand new tech industry with hundreds of billions in private capital developing commercial models. If the American private industry does actually wash out vs Chinese open models there might be talent available for them to put money into, so maybe they are just preparing for that scenario in the meantime.

MangoCoffee56 minutes ago

The American attitude is generally to let private companies build up a new industry so it can create jobs and pay taxes. However, in the LLM race, the Chinese open weight playbook pretty much killed that. China has basically commoditized LLMs. Chinese models are good enough, so the race has come down to who can offer the cheapest tokens.

yewenjie2 hours ago

I couldn't find any details about size or training data for the model.

robotbikes2 hours ago

It looks like they're taking applications for training data (due August 14th), so I think it's safe to say this is just an announcement of intent and a call for involvement vs. something that is readily available. Seems almost quaint in comparison to the strategy of sucking up every piece of data you can find anywhere on the Internet and feeding it to your LLM but I suspect their intent is to be more careful in what they train their model on.

villish42 minutes ago

I have no doubt companies like Microsoft, Amazon, and Google will rush to give them all the data they want in order to keep those government contracts flowing.

thegreatpeter2 hours ago

Pretty cool I’ll take it. Thanks!

riffic1 hour ago

stewards of the nuclear weapons biz. they'll do great here.

rozal2 hours ago

[dead]

actionfromafar2 hours ago

[flagged]

calvinmorrison2 hours ago

[flagged]

mrloopex2 hours ago

You and me both.

Triphibian2 hours ago

Sounds like a job for the U.S. Department of Shitposting

dyauspitr2 hours ago

It is. Depending on who Trump has fired or put in charge of a department it can be another shell that pumps out low quality crap. It might be the most valuable contribution on this thread.

fakeBeerDrinker2 hours ago

[flagged]

shenenee1 hour ago

Genesis is skynet

placedrock1 hour ago

Modeling with my life as data.