Back

Nvidia Nemotron 3.5 Lightning and NeMo Switchyard

96 points2 hoursblogs.nvidia.com
docheinestages16 minutes ago

They conveniently decided not to include the Qwen range of models in the Artificial Analysis graph, except the out-of-league Max variant. At least be brave and honest.

jmward0155 minutes ago

One major consequence of the ramapocalypse, I think, is an even higher focus on small efficient models. I personally believe that the multi-trillion parameter models are fundamentally missing things and the push to smaller, more efficient will drive evolutionary structural changes that will lead to future gains

schainks50 minutes ago

I am literally betting my company on this being true.

jmward0126 minutes ago

What company? I am 100% focused on this as a concept in my own internal research.

oblio25 minutes ago

It's a bad bet, historically.

I'm having an extremely hard time thinking of companies that have prospered due to optimization. Most of them were swept away by hardware advances, instead.

jmward0118 minutes ago

The 1980's US car industry comes to mind. Nearly wiped out because they refused to make efficient vehicles. SpaceX is arguably showing how a rethink towards efficient can take over an entire industry. I am sure there are strong examples in software as well but they aren't coming to mind.

I think when successful, optimization really just means 'finally built right' and people forget the ridiculously inefficient ways before.

NBJack43 minutes ago

I honestly hope to see this across all applications, games, services, operating systems, etc. We've been in a period of wasteful RAM usage for over a decade. Constraints, whatever their origin, can be a good thing.

pjmlp27 minutes ago

Same here, back to when algorithms and data structures mattered.

oblio21 minutes ago

If China makes half decent RAM I would bet more on things like 128GM of RAM being the default on low spec laptops 10 years from now.

While I do love optimized software, the hardware side, especially for PCs, has been stagnating for way too long. At least now we have a valid use case for doubling available RAM every 2-3 years again.

I had a reasonably beefy Lenovo consumer line laptop that I bought in 2011, 8GBs of RAM. Its screen hinge broke and I couldn't repair it but I'm fairly sure it was otherwise still usable in 2023-24, once the HDD was replaced with an SSD. I think even now entry level laptops are sold with 8GB of RAM.

By comparison a PC from 2000 was utterly unusable in 2012-13.

thehamkercat2 hours ago

> NeMo Switchyard, an open source library for smart routing

> When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job

How do routers like this handle prompt caching when you send the second request?

Sticky models per session? but then the second message of that session won't be sent to a suitable model, and will only be sent to the same model as previous one.

eli2 hours ago

I've seen ones that are configurable to pick a trade off point between lower cost (cache stickiness) and routing performance (best model for that turn).

But yeah I'm skeptical all this overhead is worth it.

embedding-shape2 hours ago

The repo is probably a better entrypoint to it, bit more concise description than the press releases: https://github.com/NVIDIA-NeMo/Switchyard (Notably: "Experimental software. Not for production use."). Unclear if they actually want you to deploy it or not, press release says yes, README says no, do with that what you will.

Doesn't seem to mention "cache" in the README nor the docs, but the code has mentions of it (https://github.com/search?q=repo%3ANVIDIA-NeMo%2FSwitchyard+...), I'm not sure what their thinking is there. "Good luck" essentially? Seems to be per-provider at best, but weird position for a routing library to take.

thehamkercat2 hours ago

I personally think it's snake-oil marketing with all these smart-model-routing products/projects

prompt-cache won't work with these

try-working2 hours ago

To keep it simple, forget about routers and imagine you're in Cursor using GPT for a while, reaching a cache of says 200k.

You decide to switch to DeepSeek in the same session via the model picker, and continue as usual. What happens is that the cache for DeepSeek is created with the 200k + the incremental message. After this, cache can be kept warm for both models; two instances of the cache exists, one for GPT and one for DS.

You switch back to GPT. The whole session is sent to the model with the 200k original from GPT and the incremental messages you sent to DS. The 200k is read from cache and the incrementals are new, and then added to the cache.

Let's say every second message you switch between GPT and DS; cache was 200k and each incremental message is 1k. If you kept going with only GPT, cache hit rate would be 200k/(200k+1k) = 99.5%. When you switch between two models with warm cache, hit rate instead becomes 200k/(200k+2k) = 99%.

Model routers work the same way. Keep the cache warm, replicate it in two places. For this reason, when you set up your model pool for routing, you want to keep the model pool small and differentiated.

First principles of model routing: https://try.works/first-principles-of-model-routing

role-model router and protocol: https://github.com/try-works/role-model

note: edited to keep the answer to the below message clearer

+1
hedgehog2 hours ago
+2
thehamkercat2 hours ago
average_bloke1 hour ago

I would like to propose something:

- problem: massive deluge of information because of AI

- solution: human beings should adopt a minimalist style of communicating in writing.

- e.g. this entire website page can be ten bullet points.

stavros22 minutes ago

While I agree with the spirit, I don't think the solution to bad prose is slightly less bad prose. We can write good prose instead.

encrux1 hour ago

In my opinion: the only way forward is zero-knowledge-proof authenticated social media.

We can’t have legitimate debate if we have to assume a few bad actors are cloning their voice by the thousands, poisoning debate.

If we can pin one account to a real person, we won’t get rid of LLM-content and misinformation, but at least we can hold them accountable.

ttoinou1 hour ago

Is the network based on trust and peer to peer confirmation of private keys from who you know in real life that you validated isn’t a robot ?

Or do you have something else in mind ?

kubelsmieci1 hour ago

> We can’t have legitimate debate

I'm not sure people really want that

npunt47 minutes ago

[dead]

SMAAART1 hour ago

[flagged]

marsven_4221 hour ago

[dead]

WalterGR2 hours ago

24 comments so far about Nemotron on this earlier submission: https://news.ycombinator.com/item?id=49257947

CurbStomper60 minutes ago

[dead]

XCSme2 hours ago

The new Meta 30B models seems A LOT better:

https://aibenchy.com/compare/meta-muse-glimmer-30b-xhigh/nvi...

thehamkercat2 hours ago

Muse Glimmer 30B seems to be on par with Qwen 3.6 27B (4 months old)

but

Qwen 3.8 27B is dropping this week...

XCSme1 hour ago

Yes, I was surprised to see doing it as well as Qwen 3.7 27b.

Even though that model is already "old", qwen was way ahead everyone else in that size category before this Meta model.

Also, probably for non-Chinese usage, using a non-Chinese model might lead to better results.

eli2 hours ago

The top 4 models on that site are all variants of Gemini Flash? That does not match my experience at all.

markasoftware23 minutes ago

yep, the person you're responding to created the benchmark and is using HN comments as advertisement.

XCSme2 hours ago

I should add a F.a.q. for this question.

The suite is across many categories, not only coding, and most of the tasks are low-horizon (or what the opposite of long-horizon is), where the max thinking time is around 10 minutes.

Gemini models are really smart, unfortunately they don't play well with any harness, so hard to use in practice.

But try them out for one-shot tasks, they are really good. Don't use them for coding in a harness, but you can ask them to generate code/planning (still, for coding only other models are indeed recommended).

khimaros24 minutes ago

lightning is sparse, glimmer is dense

Tactical452 hours ago

At what cost difference?

XCSme2 hours ago

I don't think it matters, if it's for local/on-device usage.

The cost is similar vram footprint I guess (?)