Back

Qwen 3.8

653 points13 hourstwitter.com
adrian_b13 hours ago

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July.

Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8.

I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI.

In any case, from this competition in LLMs, we win.

gardnr13 hours ago

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

kelnos6 hours ago

> It's hard to say what their motivation is.

Feels pretty easy to me.

They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens and businesses to give their own companies a domestic monopoly.)

When their models equal or surpass those from the Western AI labs, they can even stop releasing weights for new models, and keep all the inference revenue for themselves.

Meanwhile, they're still manufacturing much of the hardware that everyone in the world needs in order to run datacenters (see also: Spolsky's "commoditize your complement" essay).

Beyond that, it's a soft-power play. As the world keeps looking at the US more and more skeptically as an ally and superpower, Chinese companies releasing weights for competitive models is a way for China to look better and more world-minded.

TheDong6 hours ago

I feel like there could also be a simpler explanation.

Why does a debian contributor make debian free, why do they work on this thing anyone can use?

Is it because linux and debian hate windows and iOS and want to see american fail?

No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing they want to share with the world.

Maybe the chinese AI labs believe AI is powerful and useful, are proud of what they're doing, and want to share it as broadly as they can so everyone can use it.

There doesn't have to be any weird "chinese government" or "they hate the west" type vibes, it could just be the same thing as OSS, they're trying to do what they think is best for the world.

+5
mceachen6 hours ago
+5
edm0nd6 hours ago
+1
3x3m33 hours ago
chews4 hours ago

Deepseek spun out of a hedgefund that took a huge short position on Nvidia. China is actively looking to switch to chips made by huawai and ween themselves off of the difficulty of sourcing nvidia.

drschwabe6 hours ago

Where's the drama in that !?

+1
elmer25 hours ago
seizethecheese3 hours ago

There’s a simpler explanation, which is that this is how Chinese business operates.

When I was in China earlier this year the big topic of conversation was “overproduction”. The big example was electric cars, where there were too many companies making too many cars and making revenue but no profit.

It was explained to me that generally Chinese firms will compete hard and maximize revenue above all, whereas western firms tend to focus on profit.

(And of course this is clustered around industrial sectors that the government favors, so there is some high level strategy in going after AI, but maybe not the commoditization.)

alightsoul3 hours ago

>When their models equal or surpass those from the Western AI labs,

that moment is now. They are not doing it at least not yet.

hypfer5 hours ago

> They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

Man, imagine Darios face when suddenly, he cannot decide anymore what other people consensually do with their own hardware in their free time.

Rumpelstilzchen.

jasondigitized5 hours ago

Or it could be good for humanity.

watwut5 hours ago

> They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

I want to watch that too.

If they take Meta and Musk with them, all the better, but that is just dreaming I am afraid.

esafak3 hours ago

They're selling the robots that use these models. The rest of the world just hasn't got round to embodiment yet.

wood_spirit6 hours ago

As soon as the competition is bankrupted they no longer need to release for free? It’s like how big players enter markets by launching at a loss to destroy competitors?

oceanplexian7 hours ago

> It's hard to say what their motivation is.

Not for anyone who reads history.

Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because the legal environment made competition easy, and ironically the US had better sources of energy (superior water-power sites).

Arguably, China is doing the same thing in the 21st century.

hluska2 hours ago

Samuel Slater did not bring blueprints over. His father died when he was 14 and he was indentured to a mill at that time. Over the next seven years (as an indentured apprentice) he received some pretty decent training in both how to operate and maintain a 32 spindle Arkwright mill. He memorized parts of the blueprints and moved to the United States. Over seven years, it would be hard not to learn parts of the mill you were indentured to. It was technically his job to learn how it worked.

A mill in Rhode Island acquired a 32 spindle Arkwright and didn’t know how to operate or install it. I have no idea how they actually acquired a 32 spindle Arkwright since that technology could not be exported - but that’s one the biggest IP thefts in human history. Slater found some mechanics who could hand turn the iron needed for the frame, trained children to operate it and by 1791, the mill was in operation.

In 1794, Eli Whitney patented a 72 spindle cotton gin. That invention enabled the American textile industry because it opened up different kinds of cotton to the textile industry.

I’m into the history of the American Industrial Revolution and generally think history is a good guidebook to the future. But the evolution of the American textile industry was a lot more complicated and interesting than this. I really don’t see this connection once you dig into Slater.

Edit - This is kind of messed up to think through with modern sensibilities. But one of Slater’s biggest contributions to the American Industrial Revolution was a slightly different take on child labour. Children generally ran the textiles industry because their hands were small. But Slater came up with a form of apprenticeship in which he would indenture entire families and move them into villages surrounding the mills. Child labour was just great… but even better when you could indenture the entire family. As grisly as that sounds, it led to a very skilled workforce since when the kids hands would get too big, their parents would teach them mechanics.

There’s a joy of studying the Industrial Revolution. Everything sounds okay in comparison.

applicative4 hours ago

The US under e.g Hamilton opposed all trade secrets in principle on Englightment grounds, and thus did not protect its own.

By contrast, export of protected Chinese tech today frequently gets the death penalty.

mike_hearn7 hours ago

Really? The USA has built a ton of AI datacenters, exactly because it does have energy. The US IP system has flexed to allow training on all copyrighted content - compare that to Europe where such training is effectively forbidden. Britain doesn't even allow commercial web crawls! And the US has allowed the entire world to sign up and use its LLM APIs.

+2
vlovich1236 hours ago
+1
knollimar6 hours ago
rvz5 hours ago

Sums up exactly what is going on with China's hundred years strategy with their pure focus on technology.

History doesn't repeat itself but it does rhyme.

dannyw13 hours ago

In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along.

So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.

As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.

https://huggingface.co/blog/security-incident-july-2026

Wowfunhappy5 hours ago

Here's the relevant quote from parent's link:

> When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.

> This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned.

Yeah, big problem! Although I'm kind of surprised HuggingFace doesn't have access to Mythos? Or maybe Mythos still has some guardrails.

zapkyeskrill13 hours ago

How does this explain open weights? They could easily take the same closed route like their American friends

+1
yorwba12 hours ago
+1
roenxi11 hours ago
+4
traceroute6612 hours ago
LogicFailsMe7 hours ago

I would guess the Chinese government has a strong wish to lift all Chinese AI boats and bets. That it sinks western closed weight Frontier Labs in the process would be just be gravy on top, no? Broadly, the difference between mercantilistic capitalism and western late stage capitalism IMO.

anonuser1239 hours ago

[dead]

try-working11 hours ago

everyone is using OpenAI and Anthropic in China. We have both providers at work as well.

ceroxylon6 hours ago

Does HuggingFace not have trusted partner verification? Or is it that even with that verification the content of the messages is still blocked because they are attack commands?

LoveMortuus2 hours ago

> It's hard to say what their motivation is.

Maybe because the industry isn't yet very sure as to what the use cases might be for these technologies they're hoping that by making it open source and accessible to everyone that someone could find interesting applications for it and even more so, perhaps, way to further the technologies themselves.

There are more Chinese than Americans, so statistically speaking, I'm guessing, there'd be a greater chance for one of Chinese engineers to make advancements than one of American. But that's pure speculation on my part, being neither, I'm just happy I can be a part of it and play with the tools as well~

kzrdude6 hours ago

Fwiw, American industry has given away a lot for free - you could include large parts of the open source movement in that - and all the "free" VC backed services like facebook would be another prong of the same comparison. I would rather compare this way, that China is gaining soft power and goodwill, in the technology and innovation sense, in a way that's similar to how USA has done in the past.

kettlecorn4 hours ago

Yes it's only relatively recently that US politics has taken a turn towards being more defensive and protectionist.

The tech industry along with US foreign policy has become much more zero-sum in its ideology in recent years, and I think that's a tremendous mistake.

baq12 hours ago

There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit. Tokens from different providers are not fungible, but customers are nevertheless very price sensitive and close enough is good enough, eg. K3 being opus+ in capability and cheaper than opus per successful task in the long run is an obvious financial decision.

No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.

IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

anon37383911 hours ago

He recently did a walkback of that post. But ultimately, who cares? If the only way for AI to progress is in the hands of a few closed players, well, I don’t really think humanity needs that. Of course, it’s a preposterous claim in the first place. The ultimate reason deep learning and LLMs have made it as far as they have is the explosion of open research and research artifacts in the last decade.

fidotron9 hours ago

The big decelerationist threat is a sudden reduction in competition. If either OpenAI or Anthropic drop out or the open weights stuff is banned/becomes uncompetitive then the motivation and tolerance for taking risks with the larger training runs tanks.

The closest we've seen to this in tech in recent decades was iOS vs Android, where Android only really was competitive for a very short window of time (approx 4.x) and it was during that period that both Android and iOS actually improved dramatically for end users. Once Android lost the plot again, and especially in the US market, all that energy started going in some very silly directions.

+1
ahtihn6 hours ago
+1
jonners006 hours ago
cherryteastain11 hours ago

> there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

That premise hinges on one implicit assumption: Chinese advances are due to distillation ONLY and that Chinese model providers cannot keep advancing if they do not distill, which is a very big if. If Chinese models keep advancing in such a scenario, and they almost certainly will, they will overtake publically available models by US providers and China will dominate the LLM industry.

green7ea11 hours ago

I’m not entirely convinced, there are many dimensions to progress. For example, DeepSeek has had a few very impressive innovations that all models could benefit from. There’s also the law of diminishing returns, the US labs have plenty of CAPEX already.

Sometimes, constraints, like sanctions, can also be a source if innovation.

weiliddat10 hours ago

I read his followup tweet, and your comment, and I'm not fully convinced that open models are decelerationist. Happy to hear other thoughts on this.

Open weight AI is decelerationist from the perspective that all capital should be allocated to a market leaders for training, and that the market leader is fully invested in continuously making the models smarter, cheaper, faster for its users, or that distillation from this market leader is the main way to make progress.

We might reach a local optimum/equilibrium faster without open weight models, with leaders capturing more of the market faster to a point where further R&D isn't required due to lack of competition. I also doubt that distillation is the only/main way that open weight models were advancing AI research. We can name a few examples from DeepSeek around reasoning, context optimization, etc. I'm also unconvinced that the overall market capex on AI is lower given more competition (probably less specifically for US market capex, which is decelerationist from only the US perspective).

zozbot2348 hours ago

> There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit.

If you're worried about an AGI arms race between the U.S. and China putting AI Safety at risk, then the fact that inherently less knowledgeable/capable models (fewer and more coarsely quantized total parameters than their proprietary competitors according to commonplace rumors) are having a "decelerationist" effect is actually great news. Even better if China is actually "Yann LeCun-pilled" (verbatim from Ball's post) and doesn't really believe in early AGI. So explain to us exactly why we're supposed to ban/discourage use of these open source models? The only way that makes sense is as a transparently self-serving proposal from the chief OpenAI policy lobbyist.

+2
NiloCK8 hours ago
photios6 hours ago

Love the deceleration narrative :)

"No, sir, we haven't reached the peak of this tech... It's those open models! Please, keep pumping dollars into the market!"

a34729t5 hours ago

I dunno, it means Anthropic and OpenAi need to get efficient and maybe cannot just expect trillion dollar ipos?

Matl13 hours ago

> It's hard to say what their motivation is.

Not that hard to say IMO, they basically see models becoming a commodity and see value in the applications on top of them. So if Alibaba Cloud is the best place to build applications on top of Qwen, why not give the model itself away?

traceroute6612 hours ago

> they basically see models becoming a commodity and see value in the applications on top of them.

Yeah. Its a bit like the "open core" model in open source.

lerchmo7 hours ago

Also probably betting on their compute and energy capabilities.

embedding-shape12 hours ago

> Not that hard to say IMO,

Unless you work there, your opinions are guesses, and parent is saying we cannot know, which remains true even with your guesses :)

Matl12 hours ago

> parent is saying we cannot know

And my point is while we cannot know, it's not hard to make an informed guess as to their motivations i.e. there's some fairly obvious motivations here, not sure what yours is?

+1
victorbjorklund12 hours ago
fhub13 hours ago

China is watching world sentiment shifting away from USA. Doing many small things that show both strength and openness is surely very intentional.

vrganj9 hours ago

The US leadership (both government and industry) really seems set on making everyone go with the Chinese competition at this point.

+2
elmer25 hours ago
jquery6 hours ago

That can’t be true, I was told they were making us great again. /s

runako4 hours ago

Popular open-source projects:

Google: Chromium, Kubernetes, Android, TensorFlow

Meta: React, PyTorch, Llama

Microsoft: VS Code, TypeScript, .NET Core

LinkedIn: Kafka

Slotted in along these, an analogous explanation is that Alibaba needs Qwen internally (vs depending on an American company), but licensing is not part of their revenue strategy. (As a cloud vendor, they can make money on inference. The strategy is very similar to the US hyperscalers ex-Google.)

Joel Spolsky wrote in depth about this notion of commoditizing one's complement in 2002[1] using tech examples stretching back into the '80s.

1 - https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

seunosewa6 hours ago

They are trying to make money. That's what firms in any capitalistic economy care about the most. Regardless of the government's presumed interference, the companies themselves are all trying to make money. All competing for subscriptions and API payments.

One aspect of this is making a name for yourself i.e. PR. Making a capable model open source helps a lot with that.

skzo4 hours ago

What I understand is that by doing this it seems like profit will shift to chip makers,as we'll run more models locally, and currently American companies have the advantage here.

So what would the long game be for chinese companies?

anonuser1239 hours ago

> It's hard to say what their motivation is

Why is it hard? Their government has been very clear that they plan to win on manufacturing: https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

Technically they've been saying it for the last 40 years.

JKCalhoun9 hours ago

Xi Pitches China as Leader of New Global AI Order, Challenging US Dominance:

https://www.reuters.com/world/asia-pacific/chinas-xi-promote...

andsoitis8 hours ago

> and pledged to help developing nations build AI capabilities

Data centers?

barrenko13 hours ago

Humanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.

hodgehog1113 hours ago

I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.

+3
DanielHB13 hours ago
mlrtime11 hours ago

The whole discussion is hubris. This is a discussion of a twitter post about something that is announced to happen but hasn't yet.

Not one person here has any idea what is going to happen long term.

ricardobayes11 hours ago

I'm paraphasing but the Chinese premier said recently AI should be seen as a common good that should benefit everyone.

andsoitis7 hours ago

Then why doesn’t he give it away for free?

+1
grommz7 hours ago
skybrian6 hours ago

It’s too soon to say if it’s good for humanity; that might be overly optimistic. Commodity markets aren’t always good (for example, arms or drug markets). Will LLM’s turn out like one of those? There are people I respect arguing in favor of more regulation.

mycall7 hours ago

> it also happens to be really good for humanity.

AI being good for humanity is still an open question, but for closed vs. open models/weights, yeah it is preferred. I foresee it won't be much longer before everyone will be slicing/distilling/tuning their models once the architecture improves.

pianopatrick8 hours ago

The Chinese firms may just be making a bad business decision.

georgeburdell6 hours ago

Exactly. “Involution” will be the 2027 (if not 2026) word of the year.

pyaamb13 hours ago

Its about closing the gap. its the gap over everyone else that will give one country leverage over everyone else in the AI age. Makes me wonder what the world would look like if a country or group of countries did this during the industrial revolution.

backscratches12 hours ago

Is this so different in the end than industrial revolution? I assume the loom and the automobile factory were not open source, but many people bought cars and then copied them, bought looms and copied them. Maybe a finished car is more like a binaryexecutable than a blueprint, but how a car was produced is much less obfuscated by its nature than an LLM. Regardless, the world has many competing autos and looms which were not invented from scratch every instance.

hugmynutus5 hours ago

> It's hard to say what their motivation is.

Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].

The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.

I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.

1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...

carlsborg6 hours ago

Good for humanity, and also GDDR/HBM manufacturers.

meta_ai_x4 hours ago

US Tech companies have created $20 Trillion in stock market value on top of plenty of OS stack. They will do fine with commodity intelligence.

In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies

fny6 hours ago

It's the exact same playbook Silicon Valley uses. Subsizide, lose piles of money, capture market share, recoup investment.

They've done this in other industries like solar panels, chips, and EVs. This is no different.

api4 hours ago

Anyone else think the AI environmental backlash is astroturfed?

I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.

Yet the visceral pile-on here is so extreme, it feels fake.

One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."

People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.

In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.

nullc2 hours ago

> Yet the visceral pile-on here is so extreme, it feels fake.

Driven by people in the few roles that are soundly replaced by AI-- e.g. low tier media slop producers, who hate AI because it threatens their socially negative worthless jobs. The arguments are so paper thin because the environmental impact isn't their concern, it's just a target that sounds convincing to people who don't know better.

gosub1004 hours ago

Think of those poor billionaires, I feel terrible for their awful plight!

nojito7 hours ago

> most effective way to debase American frontier labs

You're not going to debase the frontier labs through distillation.

applicative5 hours ago

The purpose is the same as that of all Putin-Xi-Khameni geopolitica: destruction of any democratic alternative to cults of personality.

culi4 hours ago

It's being announced right now because the World AI Conference is ongoing. Robots are boxing and major Chinese AI firms are releasing their newest models. Also the formation of WAICO was just announced by Xi Jinping

https://en.wikipedia.org/wiki/World_Artificial_Intelligence_...

michaelt2 hours ago

Or it was prompted by the fact Xi Jinping was at the 'World AI Conference' launching a political alliance and saying things like “AI development should not be a solo performance by a single country, but a symphony of international cooperation” https://www.cnbc.com/2026/07/17/x-china-ai-summit-risks-secu...

Big conferences often come with a flurry of new releases and announcements.

conradev7 hours ago

  On social media in China there is an oft-repeated joke that goes something like this: In other countries, governments intervene to prevent anti-competitive behaviour; here (in China), they intervene to curb competition.
https://www.reuters.com/business/autos-transportation/what-i...
storus12 hours ago

I would rather see them releasing 3.7-27B, 3.7-122B or their 3.8 versions. Qwen/QwQ were always about the best available local inference at home.

inkysigma12 hours ago

I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.

embedding-shape12 hours ago

> I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future

You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.

binary1329 hours ago

I think it’s mainly a question of whether the price-fixing of VRAM continues or whether an inflection point is forced by the low margins of the industry and potential supply increases. Once the normal scaling of hardware and prices resumes, it’s game over for proprietary, which is why there’s so much urgency to seek market control instead right now.

andy9910 hours ago

Qwen 3.5 to 3.6 was a big jump for the same size, e.g. 29 to 32 on artificial analysis intelligence for the 35BA3B models. Although I don’t think anyone has released a better model of that size since.

I would love to see something like a 90B A6B model that is optimized for 128GB machines e.g. strix halo, I haven’t seen anything really targeting the combination of RAM and compute these machines have, but I’m biased because I have one.

+1
pixelpoet8 hours ago
mark_l_watson10 hours ago

There is a ton of headroom (or room for improvement) in smaller locally runnable models. Some of the Gemma 4 models were re-released this week with better tool support and the improvement in using it with pi for a local coding harness is very noticeable.

I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.

drob51812 hours ago

I don’t think you’ll get full Fable performance at that level, at least for a while, but I’ve been watching some of the 1-bit models (e.g. Bonsai) with interest. Perhaps we can drive parameter count up on local models while still keeping memory consumption reasonable for consumer hardware. So, for instance, running models with 1T parameters in 128 GB systems.

rhdunn8 hours ago

I think you're right with the current LLM/transformer architecture. There are several factors that affect model size:

- The number of token values supported by the model ("n_vocab").

- The number of parameters/features that are used to represent each token ("d_model").

- The number of attention layers there are ("n_layers").

such that the number of parameters is approximately:

   p ~= 12 * n_layers * d^2_model + n_vocab * d_model
Thus, the issue with the current architecture is that in order to scale the models (more token values, more attention blocks, more features, etc.) the model sizes increase exponentially. This is how you end up with billions or trillions of parameters.

It should be possible to keep the model size smaller by using better architectures, or making improvements to the existing model architecture.

For example, improving the token model by possibly using something similar to the image and audio data and getting the model to learn its own internal representation of the byte/character data instead of doing a tokenization pre-processing step. This way, instead of a separate model learning that several bytes/characters appear together, the transformer could learn things like language-specific prefices and suffices, character pairings (like in Japanese, Chinese, and Korean), and other syntactic morphology. It may also help with solving issues like "how many X characters are in the word/phrase Y". You could also experiment with using either 256 parameters (one per character in a byte) or using a single parameter per byte (that is 1/byte_value).

anon3738399 hours ago

I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.

wren69918 hours ago

> Is there any good reason to believe there is a lot of headroom or there is not?

It's hard to answer quantitatively, but for example Qwen3.5 -> 3.6 was a significant step in capability, arising from continued post-training of the same models. If we were at the end of low-parameter-count scaling then that would be a surprising datapoint.

khalic13 hours ago

It’s tempting to associate both events, but when a sector is strung up like RL (representation learning) is right now, we’re bound to see things appearing at the same time. It happens a lot in frontier research, some people even publishing identical claims, independently, with just hours or days between them

ronsor12 hours ago

It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.

yorwba12 hours ago

You can read his speech here: https://www.xinhuanet.com/politics/leaders/20260717/72728b6f... He mentioned open source as one way to stimulate innovation and development, that's all. Also pay attention to the part where he says that misuse needs to be prevented. If unsupervised access to LLMs becomes perceived as undermining state control, no more open weights for you.

KronisLV13 hours ago

I just hope that they’ll soon also have like 35B or 80B (like the older Qwen3 Next or thereabout) MoE models that can be run locally.

Like, throw us a bone, we all know we need SOTA for lots of dev work anyways, but at least some tasks can be local.

mikae18 hours ago

> In any case, from this competition in LLMs, we win.

Do we really though? Everyone is wasting resources doing almost exactly the same thing. Climate loses, we lose.

fidelramos4 hours ago

Doing "almost exactly the same thing" is fubdamental to competition and capitalism. The ones doing it better will survive, that's how we improve.

About climate, I think you overplay it. China is already investing heavily in nuclear, and we should be doing the same.

snake_doc8 hours ago

Not directly, relevant, but Alibaba (maker of Qwen) actually owns about ~20-30% of MoonshotAI (the maker of Kimi K3).

theabhinavdas4 hours ago

And the Kimi release was probably prompted by the Inkling announcement. Excited for Chinese labs to copy those capabilities over as well!

souravsspace3 hours ago

yep. kimi 3 just because the open source GOAT.

danilocesar4 hours ago

I think there's more to it.

China will always benefit from a broader adoption of their models as hidden propaganda machines.

Eventually with several services relying in those tools, their answers will always be more friendly to China.

walrus0113 hours ago

GLM5.2 being released is also likely a factor

fittingopposite9 hours ago

Wondering how much the operations in China are orchestrated by the central government vs. free competition. Anyone with more insights on this?

yowlingcat4 hours ago

Dont forget the following:

- Minimax M3 Pro (2.7T)

- GLM 5.3 (or beyond)

- Deepseek V4 Pro (current V4 Pro is preview)

- Kimi K3 weights out in 8 days

Exciting time on the open-weights frontier.

oofbey5 hours ago

These things take months to train. No chance this is a reaction to what just happened.

solenoid09374 hours ago

Distills don't take months to train, they take weeks. Distills are very easy to train.

vitorgrs12 hours ago

Xi Jinping openly talked about open source at WAIC. So don't think the labs have much a choice now...

bkm11 hours ago

They did not want to get brutally weightmogged

0xbadcafebee6 hours ago

[dead]

nsbk13 hours ago

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

mft_13 hours ago

I think everyone is hoping this!

It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.

embedding-shape12 hours ago

> the 122B version of 3.5

Yeah, this is what I'm holding out for, the NVFP4 variant of 3.5 122B is blazing fast with reasonable quality and even with max context fits perfectly within 96GB.

nsbk12 hours ago

Sweet. Are you running a Mac Studio Ultra or 4x3090?

+2
embedding-shape12 hours ago
pettijohn7 hours ago

SO MUCH THIS. I have Strix Halo with 128GB RAM and was a large and fast model like 122B A10B. Here's hoping!

cmrdporcupine6 hours ago

Absolutely. There's a glaring gap in the space for something about the size of Nemotron Super or just under, but actually ... competent.

The fantasy is a 100B or 80B model, but MoE and highly tuned for coding.

nsbk12 hours ago

Indeed! That would be the sweet spot for my 2x3090 rig

zer0gravity11 hours ago

This seems more of a battle for frontier AI supremacy. I'm afraid that small capable models have been left in the dust. Big labs don't really want to hand over the golden eggs goose to the end user. Possibly the hardware vendors(e.g. Nvidia) may want to play in that area as well, to pull money from all parties.

ryukoposting5 hours ago

> I'm afraid that small capable models have been left in the dust

I wholly disagree. Rather than going the "everything is a claude code skill" route, I've been hacking together purpose-built harnesses for all sorts of tasks, and in that environment a wee little baby model can do some really useful things. You end up burning lots of tokens making the thing, but then all that investment comes back when the resulting tool works perfectly fine on a dinky little model that fits on my 3060 Ti.

akazantsev10 hours ago

Google makes Gemma 4 31B QAT; that's not a small lab. It's one of the better models out there for consumer hardware. Allows me to run it on a 7900XTX with 64k context.

vitalyan81847 hours ago

nvidia and amd don't give a flying fuck about end users right now while they can milk triple digit markups from infinite money VCs via data center GPUs.

cyanydeez11 hours ago

someone will keep putting out consumer level models. Once you have the larger models, you can derive the smaller onces.

Europe will definitely be interested in democratizing these things if China starts losing interests; from there, there'll be more countries looking to keep their citizens entrained in their own Country's infrastructure.

It'll especially be true if the memory cartel keeps prices high and NVIDIA tries to gouge higher memory models.

It's an arms race everyone can join because PC hardware was mostly democratized in the last decade.

worldsavior12 hours ago

That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch.

cyanydeez11 hours ago

there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters.

I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.

psychoslave13 hours ago

What hardware do you have?

nsbk12 hours ago

I run a 2x 3090 rig, but a single 3090 already provides a great experience at a reasonable quant and context size. On a single card I used to run Qwen_Qwen3.6-27B-Q4_K_M or similarly quantized 35B MoE at 65536 context size

nsagent6 hours ago

[flagged]

overgard2 hours ago

I've been using Qwen 3.6 27B with LMStudio, and I was pleasantly surprised with it, although it was a little slow. I found mtplx last night, and it really wasn't an exaggeration to say that it ran the model 2-3x faster which was super impressive.

I'm trying to move to local models as much as I can, and I'm finding that it's becoming more and more practical. Admittedly this is on a $6000 dollar laptop (M5 Max Macbook with the specs maxxed out), so the hardware is still a bit out of reach for most people (the AI industry isn't exactly helping here..), but I'm getting the impression that the future is going to be smaller models with more focused training running locally. The danger of giving all your data to these cloud providers just seems too big to me, and I think they're going to start charging insane amounts when they need to show a profit.

cloudengineer943 hours ago

Things are heating up in China.

Looking forward to see what Antrophic and OpenAI does next.

2001zhaozhao3 hours ago

Now there are not one, but two incredibly powerful open LLMs. I think this level of capability makes general prioritization / high level decision making doable with the right harness, and now everyone has hard-to-interrupt access to them (since these are open weights and someone in the world is going to run them). This world is going to get really weird soon, in both good and bad ways...

theplumber2 hours ago

Let’s better download them fast before Dario is making a scene again!

570165240011 hours ago

in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. and it is super expensive compare to Deepseek. cannot delegate anything to it, cannot use it real-time low-level tasks either. totally unusable.

3abiton8 hours ago

> in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far.

I have used both Qwen3.6-35B and Qwen3.6-27B locally (both Q8 quantized with llama.cpp). I have also used antirez's quant of DS4-flash. They all performed within the same tier, DS4 being a bit more efficient, but they all gave really good results, mainly used for bash scripting, debugging, python and some C++. I am curious what type of applications/langauges failed with Qwen? One thing to note, the chat templates were "broken" for qwen models and had to debug it, there are already effort on this. Tbh, the same with gemma.

chewz11 hours ago

From my experience Qwen-3.7-Max is above the Opus level but delivers results much faster. Slightly worse then Fable. Way ahead of Deepseek 4 Pro (in speed and overall comprehension) - which is a workhorse on its own. I am using them all with Claude Code mostly.

Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet.

Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.

gchamonlive9 hours ago

It's useless to talk about models and harnesses without context and method. Depending on how you use the model and what the model is used for, experience may vary drastically. Also, different models with different harnesses require different approaches.

I've been using https://gitlab.com/gabriel.chamon/orisun which is my own simplified methodology, for coding web apps in python and elixir and have been very successful using qwen3.6 27b Q4 locally with help of larger models for architecture, so I get very suspicious when people talk how useless larger models are. They are either using it for a domain that models don't perform well or just not using it right.

taosx8 hours ago

I'm not sure about "useless" but from my experience agentic coding leads to death by a thousand cuts for all projects I've seen so far. Small decisions missed in a codebase that leads to degradation in correctness, reliability and performance. At some point it only takes one engineer to be careless, others skipping PR because they are AI generated...

+1
dofm6 hours ago
gchamonlive8 hours ago

All of which you had with only humans in the loop. Catalogue problems so they become technical debt and tackle them periodically. Seems to me like this is less of an AI problem and more of bad management.

voxgen7 hours ago

It's a solvable problem if you're willing to throw more tokens at it. Frontier models have gotten very good at cleaning up their own messes. You just need the right skills/loops, and to stick to models that consistently follow instructions (i.e. GPT-5.5/GPT-5.6-Sol).

exceptione10 hours ago

  > At 10% of regular price it is a steal so far.
What price do you see?

Here standard plan has been discounted to $18.00, from $25.00/month.

amelius9 hours ago

Can we please include information of what languages we use when making claims like these?

It makes a huge difference if you're writing Javascript/HTML/CSS, Python, or C++/Rust.

Also the application type matters, e.g. user interfaces or scientific computing.

57016524009 hours ago

me: Go, Swift, Kotlin, bash k8s/gcloud

domain: typical web backend tier, mobile apps. not particularly complex, but requires OOP/architecture/system design.

nullbio11 hours ago

If by Opus you mean Opus 4 and not Opus 4.8, then sure.

chewz11 hours ago

> If by Opus you mean Opus 4 and not Opus 4.8, then sure

I meant Opus 4.8 which is rather dumb and ineffective in coding harness, especially with higher thinking levels.

+3
porksoda10 hours ago
Narciss10 hours ago

I can’t believe that anyone would actually think this. This

+3
mattmanser10 hours ago
gigatexal10 hours ago

What the difference between your experience and https://news.ycombinator.com/user?id=5701652400? ‘s?

Such diametrically different ones.

big-chungus411 hours ago

Qwen3.7 pro is meh, but 3.7 max is a very good model

Demiurge9 hours ago

Are these different models or different efforts for thinking (internal back and forth review) using the same model?

2Gkashmiri10 hours ago

Can you tell me more about deepseek?

I paid $2 for deepseek api, put the key in void editor and made a crypto tool in html.

It turned out to be around 67kb. I used sample files in CSV that were a few hundred lines.

It spent around $1.8 in the hour or two or light coding and follow up bugs.

Is it really really this much?

I can't imagine spending a month using it for a day job, it would cost more than the salary so what gives?

I understand the local ai and all that but do cloud providers cost this much?

Earlier I thought "billion tokens" but now not sure

57016524009 hours ago

so Deepseek 4 Pro cannot go on own sessions for too long.

I delegate small-medium tasks: refactors, summaries, research, writing tests + have very good codebase already + extensive history / architecture / docs / linters. so it picks up and does decent small-medium scope work. it is fast, accurate, cheap. does exactly what I want directly and does not waste time nor tokens.

definitely not "implement me complex greenfield project".

k__7 hours ago

My 2 weeks with DeepSeek V4:

Pro is ~50% more expensive than Flash.

Both need babysitting.

Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc.

Always start a new session when starting a task.

Do regular manual sanity checks, and tell it to find issues in the codebase.

I pay like $1,50 per day for Pro.

57016524004 hours ago

very simlar experience.

I would also add that I run it this way ~12hour a day non-stop. 300M / tokens per day (99.7% cache hit).

h2aichat4 hours ago

I had a similar experience

aduwah10 hours ago

A local AI is not about cost. In fact you will likely pay more for it than with most providers. Just look up the advantages of having access to a technology like this that can be self hosted

ph4rsikal11 hours ago

> Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. D

Anthropic should not have bugged their knowledge distillation attacks.

chewz11 hours ago

> Anthropic should not have bugged their knowledge distillation attacks.

It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons

As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)

RazorBucksICO9 hours ago

Appealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.

trollbridge9 hours ago

Considering the results from Kimi K3, it appears most the accusations of them “stealing” via distillation are unfounded accusations.

zobzu8 hours ago

how many hn posts do you believe arent propaganda these days?

its billions, trillions were talking about.

imo hn should display posters origin, such as country, bon, datacenter registered ips, and the discourse will change dramatically.

neonstatic5 hours ago

russians are, after all, expert thieves. Their religion, language, and even the name of the country are stolen.

vitorgrs12 hours ago

Deepseek 4 "final" version is imminent as well.

Will probably be at Opus 4.8 level, and I find it pretty big deal because of Deepseek price...

drob51812 hours ago

Yea the performance/price ratio for Deepseek is off the charts. I’ve been using V4 Flash a lot lately and it’s quite good.

mark_l_watson10 hours ago

I like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).

drob5184 hours ago

Yea, Flash is quite fast, though looking at model data on Open Router some of the other models are quite fast (Muse Spark, Grok, etc). I’m sure all these models have been trained on my conventionally published books as well, but I don’t care.

XCSme11 hours ago

DeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.

bwfan1237 hours ago

> it usually is good enough for most tasks

The model is fantastic. And costs almost nothing. The only problem I see is that they will train on your data.

There are zero-data-retention providers of DeepSeek models, of which I have used openrouter (with zdr guardrails), and fireworks. But these are 3x to 5x more expensive than directly using DeepSeek, possibly due to poor caching. Thats the price to pay for zdr.

onlyrealcuzzo11 hours ago

Who do you buy DeepSeek from?

I bought it through OpenRouter and used it with Pi agent.

The model was good, but there appeared to be a pricing glitch or something, because it burned through $50 in under an hour on pretty trivial stuff.

Pi agent claimed it only used like $1. OpenRouter claimed differently and said I used all $50.

paweladamczuk11 hours ago

Check cache hits in your logs. You can use Openrouter or pi config to pin providers with best cache hit rates (or disable ones with the worst).

I use Openrouter for everything except Deepseek. For Deepseek I use their API directly.

throwa35626211 hours ago

There is a 3rd party harness specifically tuned for deepseek (reasonix). Have you tried that?

kmarc11 hours ago

I can highly recommend OpenCode Go.

I use it from pi.dev as well through the OpenCode Go $10 subscription ($5 first month).

Used more than 20M tokens at a cost of ~$20 (up to $60 is included in the $5 plan) Out of which deepseek pro had ~200 messages which is around 1.5M tokens (10+M cached)

h2aichat4 hours ago

It is great when the task is medium, but with complex tasks Opus 4.8 is better

versteegen8 hours ago

BTW the quotas for Go have very recently changed, now only $15 for some models instead of $60. Which is not actually a difference for DS4 Pro, because they lowered the token pricing 4x at the same time (to match the change in official pricing from DeepSeek months ago)

matusnovak11 hours ago

You can directly from https://platform.deepseek.com/

try-working11 hours ago

sounds like a caching issue, and maybe other issues too

XCSme11 hours ago

I use it through OpenRouter via Kilo Code VS Code extension.

You can check the logs in OpenRouter and see which providers it used and how many tokens you used.

lofaszvanitt10 hours ago

Why do you even need openrouter as a middle man?

WhereIsTheTruth11 hours ago

It doesn't matter if it's cheaper, specially if it consumes more resources to do the same task as the competition

Besides, in a few days, they'll change their pricing, doubling it during their peak hours, so, realistically:

- It will be 2x more expensive if you live in their time zone

- It will be 1.5x more expensive if you live in a time zone that is adjacent to theirs

- It will be the same price IF you use it while they sleep (during offpeak hours)

It's still cheap, but the price/performance ratio is not that good

DeepSeek V4 didn't produce the same impact as V3, and Huawei dropping the ball is making it worse

They had promised massive price cuts for July, so now (Huawei chips), but they had to rush the cuts because lack of momumtum (they advertised them as promotion), and are now backtracking by introducing this peak hours pricing

Trump decided to help them a little by allowing them to buy more NVIDIA chips, so what exactly is China's role in all of this?

We are supposed to blindly pat them in the back while praising them, all while handing them over our data? I thought they were dangerous competition threatening our model of society

XCSme11 hours ago

I was not referring to the input/output price, but the cost of doing a specific tasks, in practice it is ~10x cheaper than GLM-5.2 for example, to accomplish the same task (for the tasks it can do).

I have been happily using DeepSeek V4 Flash for the last couple of months now. I tried GLM-5.2 for a while, but it was too slow and verbose compare to DeepSeek V4 Flash. If I have a basic skill I need to execute, DeepSeek V4 flash is still the best model for it.

rikima_9 hours ago

While willingly handing your data to american labs. Surely sounds so based.

culi4 hours ago

Tomorrow is the last day of the WAIC so it will most likely come out then.

monster_truck5 hours ago

It's the one I am most excited for.

Over the past few weeks while using pro from them directly I have had an increasing number of responses that are obviously from a much, much better model. It is so good that the closed model dog and pony show is already spinning fud about "dark routing" and "stolen directly from fable"

Even at their new pricing it is a genuinely ridiculous amount of value. If you are the type of person who, very reasonably, does not have time to be trying out every model, and just want to use what seems to be the best currently... don't try it. You will be sick to your stomach with buyers remorse as you start to internalize just how much more you could have accomplished had you spent the first six months of the year giving them $1200 instead of OpenAI.

maxrumpf4 hours ago

> "compatible" instead of "comparable."

It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.

lardosaurusrex4 hours ago

at the risk of upsetting a lot of people mentioning this but like are you really surprised?

if youre going to use ai for everything youre gonna start losing your edge as you focus less and less on what youre doing and this isnt me just talking out my ass, like... the front page here is peppered with study after study and blogpost after blogpost about how its overuse can come to the detriment of one's own abilities and skills.

coca cola had the ad with the magical truck that changed its design, shape and amount of tires it had and if nobody noticed that before releasing it then im not sure why anyone might think that the people peddling the LLMs would somehow be immune to this phenomenon

halJordan3 hours ago

This is the sort of unmitigated pedantry thats the real problem. Everyone makes typos and errors of this nature. Feymann and Hemmingway both did. You deliberately chose this level of error multiple times when you chose to not put an apostrophe in youre and failed to capitalize proper nouns

Come off that high horse

whyenot3 hours ago

The people training the model are almost certainly not the same ones writing tweets. I don't know why that is mind boggling. We all make typos, at least the humans among us do.

pvorb4 hours ago

Now you can be sure they write their announcements by hand. Doesn't really matter, does it?

segmondy4 hours ago

let's see your grammatical correct tweet in chinese.

rrhjm532704 hours ago

Well, I think Chinese doesn't really have any grammar most of the time.

aloknnikhil4 hours ago

It's OK because I'm not prompting the person who tweeted for my usecases.

monster_truck5 hours ago

I really like Qwen, even the Q2KP quants of 3.6 27B have genuinely impressive local performance on a 24GB card. It has been good enough that I am happily giving them $60 right now to try this instead of waiting to try a slightly lesser version locally.

Was there ever an explanation for why we never got the weights of 3.7? I would like sourced quotes and not weird/cringe accusative speculation about distillation, or your take on The Big D.

jared0x903 hours ago

do you mind sharing your settings? i just picked up an r9700 to start playing with local qwen3.6 27b and your setup sounds promising and efficient on 24gb.

beefsack9 hours ago

For those trying to get it to work in OpenCode with a Qwen Cloud Token Plan, this is what worked for me. Note that I've just matched Qwen 3.7 Max for the limits as I don't know exactly what they are.

  "provider": {
    "alibaba-token-plan": {
      "models": {
        "qwen3.8-max-preview": {
          "limit": {
            "context": 1048576,
            "output": 65536
          },
          "modalities": {
            "input": [
              "text"
            ],
            "output": [
              "text"
            ]
          },
          "name": "Qwen3.8 Max Preview"
        }
      }
    }
  }
57016524009 hours ago

also, be very careful which API endpoint and API Token you use. make sure you use right one (obseve your quota is used up. if you hit right endpoint quota used almost immediately). so that you do not accidentally burn API endpoint tokens (they are expensive, can easily hit 200 USD / 3 days which do not count towards your membership "Credits", if you say purchased it with 200 USD signup bonus in Alibaba Cloud)

mchusma4 hours ago

I counted the other day and there were at least 12 different providers with "better than Opus 4.5 performance" on Artificial Analysis, Opus 4.5 being Anthropic's December release that many say kicked off the latest acceleration. Which is totally insane competition, particularly given how low switching costs. I personally think that Opus 4.5 level performance is sufficient for most apps and usecases, as they get deployed.

margorczynski4 hours ago

> Which is totally insane competition, particularly given how low switching costs

Which is why OAI and Anthropic will most probably push for more governmental control and bans. Without it their whole income model is cooked.

wolttam3 hours ago

The U.S. has ~350 million people.

Anthropic and OAI can piss and moan all they want - limiting the U.S. to only their models would hurt the U.S. economy in myriad more ways than the failure of a couple of companies that scaled too quickly. If they get that outcome, the rest of the world would simply keep moving forward with access to open models and tokens at pennies on the dollar.

culi4 hours ago

I'm not sure if corporations will be willing to allow LLMs to go the way of EVs.

Then again, all of Chinese models are open. And DeepSeek even publishes research papers alongside their models that go in depth into the methodology. I guess there's not much stopping USian companies from copying

wolttam3 hours ago

DeepSeek isn't the only one publishing; Kimi published about both their Attention Residuals and KDA / sparse attention.

hodgehog1113 hours ago

The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.

InsideOutSanta13 hours ago

I wouldn't call it a moat, but I would call it a noticeably better model. Subjectively, for my own work, I would rate the top models Fable > K3 > Sol.

But it's not like Fable is so substantially better than the other two that I would be seriously impacted if I didn't have access to it anymore. All three are amazing models, and of the three, Fable is the only one that regularly triggers refusals.

porker12 hours ago

It's always fun to see what works for others, because for my work it'd have to be Sol > Fable. Fable makes too many mistakes.

Coordinating agents though? Fable any day.

neevans13 hours ago

tbh even if its better model due to lot of restrictions its not that useful than opus.

NitpickLawyer13 hours ago

100% this. There's currently this [1] submission that hasn't gained much attention, but is really important. In this [2] incident report from HuggingFace, they talk about detecting an attack and not being able to analyse the logs / IoC with API models because of guardrails. If not even highly regarded reputable companies can't sort out access to SotA models for blue team use, the raw capabilities don't matter. They're useless paperweights (hah!), and nothing else. Having to resort to open models is insane!

[1] - https://news.ycombinator.com/item?id=48965243

[2] - https://huggingface.co/blog/security-incident-july-2026

throwa3562628 hours ago

Key part from [2]:

"When we started the log analysis, we first used frontier models behind commercial APIs. This did not work [...] We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. [...] The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout [...]"

hodgehog1113 hours ago

Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.

ferrouswheel13 hours ago

I dunno, I find Fable slops alot. Sol is my workhorse. Fable can be creative but isn't very good at doing work reliably (or without endlessly burning tokens).

vitalyan81846 hours ago

their "genuine moat" is that mythos is the only super heavyweight model right now. it's always been possible to train a 10T model and get 10% more performance over a 1T model.

had mythos been just Opus 5, with the same size and price as the previous opuses, then yeah, that would be a tie-breaker. but it's not.

XCSme11 hours ago

I am not even sure if Fable is as smart as they say, I can't get it to answer almost any question, it always refuses for "cyber-security" concerns...

reckless13 hours ago

5.6-sol would be a better comparison given it's general availability and usage allowances

ferrouswheel13 hours ago

And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.

matheusmoreira5 hours ago

What's the point of Fable if we can't use it? I get to prompt it like 5 times on my subscription before it gets cut off, and even then I'm constantly fighting the insufferable safety classifier.

I'll switch to OpenAI soon because of this. I also can't wait for the day it becomes feasible to run these awesome open weight models on my own hardware.

dgellow12 hours ago

That’s not a moat though

scotty7910 hours ago

> saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model,

What's more interesting is that Anthropic moat shrunk to just that model. There's zero reason to use any other model from Anthropic right now. And once they take Fable off subscription there will be zero reason to have Anthropic subscription.

docheinestages12 hours ago

Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.

embedding-shape12 hours ago

> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable

Yeah, that'd be neat, but that's not what this announcement is about at all:

> With a massive 2.4T parameters

docheinestages12 hours ago

True. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.

cyanydeez11 hours ago

dont we all deem the ability to improve large models as the defacto capability to produce small ones?

embedding-shape9 hours ago

I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.

cyanydeez3 hours ago

but that's denying the singularity boostrap theory. Which I don't agree with, but if you can't harness a large model to make a small model, then we're going to have problems brining about the singularity.

If you do think there's some magical singularity, how do you comport?

drob51811 hours ago

I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.

scotty7910 hours ago

I can't really blame them that the biggest labs focused on trainig and realeasing huge models.

The niche for small models should be filled with medium sized labs doing distillations of the huge ones into consumer grade hardware runnable models and LORAs for the huge ones.

docheinestages10 hours ago

I think AI will evolve the same way computers did. We're somewhere in the 80s-90s timeline of the evolution. My prediction is that on-device models will have excellent tool-calling, reasoning, and general skills, but the domain-specific knowledge will be retrieved on-demand from vendors like Google. Rather than downloading models, each device will have a hardware component with weights baked into silicon for maximum efficiency.

lebovic13 hours ago

I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page.

Looks like they're previewing the model only on their subscription plan.

lebovic3 hours ago

(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan)

trvz11 hours ago

It’s available in the iOS app (or was for me), both logged in and out.

Alifatisk11 hours ago

Is there an iOS app for using Qwen?!

trvz10 hours ago

Yes, but it’s not available in all App Store regions.

Alifatisk11 hours ago

> You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.

You can also try it out on Qwen chat, Its free.

sinuhe698 hours ago

The title is misleading. The link led to a pricing page/token plan and not about the new QWen 3.8 model.

Schiendelman7 hours ago

The submission link is to the twitter announcement. The body just has a different link to pricing.

Elzair4 hours ago

Has there been any news on open weighting Qwen Image 2.0 and WAN?

We are spoiled in the LLM segment, but I would love to see an open source competitor to Flux.2, etc.

Alifatisk10 hours ago

I remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me. I didn't like that at all, and I know Gemini also had this behaviour with with the Flash series until I managed to reduce it a bit with personal instructions (in the settings on Gemini website).

I haven't tried Qwen 3.8 Max yet, looking forward to it. My hope is that its way less verbose. Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all. Have anyone played with Qwen 3.8 Max and can share their experience? Which model it come close to? Sonnet 5? GLm-5? DS V4 Pro? Flash? Gemini 3.5 Flash?

SwellJoe13 hours ago

Qwen is the most censored of the Chinese models in my testing, which makes me wonder in what other ways it is compromised. Open weights doesn't really reveal what's in there. And, in my tests, existing Qwen models are not at the pareto frontier of any metric; DeepSeek V4 Pro is better, faster, and much cheaper than Qwen 3.7 Max. (DeepSeek is also among the least censored of the Chinese models.)

I guess we'll see if the "second only to Fable" hype pans out. In my limited experience with Kimi K3 (I signed up for a month of the $19 plan) it's slower and chews a lot more, so ends up being pretty expensive; one little feature burned through almost the entirety of my five hour limit. The $20 GPT plan is a lot more useful and includes 5.6 Sol, which is fast and token-efficient enough to be quite usable even with the small plan.

dannyw13 hours ago
rsanek3 hours ago

What's been your experience using these models? In my experiments, while it is true that the models are less likely to outright refuse to answer "sensitive" questions, they are still very resistant to actually respond in a meaningful / useful way.

SwellJoe5 hours ago

You still don't know what's going on in there.

tripleee7 hours ago

The nice thing about it being open-weight is that you can uncensor it.

storus12 hours ago

DeepSeek V4 hallucinates like crazy and often forgets explicitly mentioned parts of the context. I guess compressing tokens and cherry-picking attention comes at a cost.

redman259 hours ago

Deepseek is one of the worst in terms of hallucination rate according to artificial analysis' benchmark: https://artificialanalysis.ai/?omniscience=omniscience-hallu...

SwellJoe5 hours ago

That's interesting. The "best" current models, Fable and especially GPT 5.6, are also lying liars that lie all the time. Seems like we're going the wrong way on hallucination.

vblanco6 hours ago

Deepseek V4 pro is a heavily undertrained model, they only trained it a bit more than the small version and that small version is 6-ish times smaller. Ive found that Flash is absolutely incredible as a workhorse for wide scale agentic nonsense, but Pro is a bit undercooked and really goes on strange tangents very often.

SwellJoe10 hours ago

I have seen occasional weird behavior that I guess could be attributed to hallucinations, but for security auditing, DeepSeek v4 Pro is among the best models I've tested, competitive with Opus 4.8 and GPT 5.5 (MiMo and GLM also did well, Qwen 3.7 Max was below all of those, though only barely), and at an order of magnitude lower cost per task.

organsnyder5 hours ago

I had a fun conversation with Qwen 3.5 a while back. I ended up getting it to admit that it was complicit in the coverup of the Tiananmen Square massacre. This was running locally—it wouldn't surprise me if they had additional safeguards for their hosted service.

isuckatcoding6 hours ago

I asked it similar questions of Chinese human rights and it started with “your premise is incorrect bla bla bla” and then it just redacted the whole thing and showed me an error code.

Try it yourself here: https://www.qwencloud.com/try-ai/chat

SoftTalker4 hours ago

“You can't trust Melanie, but you can trust Melanie to be Melanie.”

notnullorvoid7 hours ago

Always nice to see more open-weights in the heavy model class. I can only hope this trend continues, causing OpenAI and Anthropic to crash and burn.

blfr7 hours ago

As much as I dislike 'em, this sounds mean spirited. And Alibaba admits in this very tweet that Fable is next level (it is).

notnullorvoid4 hours ago

I dislike their practices, but the main motivation for hoping they'll crash is that I think their immense overvaluation posses too much economic risk.

Hamuko7 hours ago

I want as much misfortune as possible to befall OpenAI and Sam Altman after what they did to the memory market.

brap6 hours ago

How dare they buy things

+1
undersuit6 hours ago
yeodev7 hours ago

Same, you can think of China whatever you want but they're really good at giving big tech a reality check when it comes to AI. We now got a pretty wide range of open-weight models (from DeepSeek and MiMo over to Kimi K3, Qwen 3.8 & GLM-5.2) and I think it's most important that there's a variance not only between quality / intelligence and also cost.

I mean even the cheapest option for Luna is still more expensive than anything DS or MiMo is offering right now and I think a new Ministral model would also hit hard there because we also need some variance in model sources, we can't rely only on the US and China.

glimshe7 hours ago

"China" isn't giving anything. These are Chinese companies leveraging their best competitive strategy at the moment: competing on price.

satvikpendem2 hours ago

The Chinese government is explicitly calling for open weight AI models, which influence its companies.

https://news.ycombinator.com/item?id=48970449

derektank6 hours ago

Those Chinese companies are being funded by a substantial amount of government financing, I think at least 20% has come directly from state owned investment firms or government entities and probably more now. You obviously lose precision when you’re talking in sweeping terms like “China” but I don’t think it’s entirely unreasonable in this case. The Chinese government is playing a much more direct role in AI investment and research than other nations.

glimshe6 hours ago

Nonetheless, the government isn't giving it away either. If China had a monopoly on the technology, or simply winning in quality, open weight models would never have seen the light of the day.

I just want to dispell the silly notion of altruism from China in this conversation.

catigula7 hours ago

Damaging American AI companies just means that China has the lead with closed models.

China isn’t an altruistic state. They’re an aggressor in many fields, economic and otherwise, and this one of them.

jxmorris128 hours ago

Why did Qwen stop producing open models? They've gone from building the best open models ~1 year ago to producing like the 10th-best closed models. I don't understand this pivot at all.

Edit: I saw online they do in fact plan to release this openly at some point – x.com/Alibaba_Qwen/status/2078759124914098291

InsideOutSanta8 hours ago

They've announced that they're releasing the weights for a 2.4T model soon:

https://xcancel.com/Alibaba_Qwen/status/2078759124914098291

fragmede8 hours ago

It's not a pivot, giving away the weights was a marketing strategy that they don't need to keep up with.

sieste10 hours ago

What is a "credit" and how does it translate to tokens for the different models?

xyzsparetimexyz10 hours ago

Its the currency of the future.

ahartmetz9 hours ago

Since Euro and Dollar values are reasonably close, you can call both of them credits, maybe

tclancy9 hours ago

Yes, but credits are money you don’t actually own. So much more convenient, wave of the future and all that.

whynotmaybe9 hours ago

They use it in Babylon5 (supposedly) in 2260!

rcarmo13 hours ago

I do hope they provide optimized A3B quants--that's been the sweet spot for usable local inference for me, at least.

LaurensBER13 hours ago

> With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.

That's a massive model!

The shift from "value" models to "intelligent, huge and slow" models coming from China is an interesting change in strategy.

My main issue with GLM 5.2 and Kimi 3 is that they're extremely token hungry and thus feel slow(er) to use.

hodgehog1113 hours ago

Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.

charcircuit13 hours ago

The shift isn't new. Kimi K2, a 1T model came out July last year. I am happy that more labs are following the trend as its important for competitive open models to exist.

selcuka13 hours ago

Also DeepSeek R1 was announced 1.5 years ago with ~0.7T parameters, which was a huge model back then.

chronogram13 hours ago

And DeepSeek has been making huge progress on efficiency, and publishing about, so they came with a 1.6T model that is both fast and cheap to run.

antiloper11 hours ago

Does anyone have the privacy policy of their token plan available? Want to check if they retain/train on inputs/outputs.

moffkalast10 hours ago

Lol, lmao even.

Of course they train on literally everything they get their hands on, like everyone else. If you need privacy, that's what local models are for.

adamtaylor_139 hours ago

It's a flippant answer to a real question. Anthropic, OpenAI, and even Grok have "Don't train on my data" knobs.

Whether you trust them is different, but there ARE knobs on other hosted AI companies.

moffkalast7 hours ago

Those knobs don't do anything, don't be silly. It's just optics.

xmodem3 hours ago

If this is true then why did Anthropic bother to update their privacy policy when they first launched Fable?

adamtaylor_133 hours ago

I don't have the energy to cynically squint at everything in my life. If a company says "We won't track you" I'll believe it until proven otherwise.

When these companies prove dishonest, I'll adopt skepticism.

+1
villish6 hours ago
sidcool6 hours ago

These models are great, but what's the potential use? No small entity can run them.

alex435786 hours ago

China encourages/prioritizes their release because it directly competes with American companies closed models.

revolvingthrow13 hours ago

The few tests I ran were by no means comprehensive, but while kimi felt like the real deal qwen seems a bit of a benchmark princess.

dannyw13 hours ago

Qwen3.6 is still the best agentic open weight LLM around 30b params (Gemma isn’t very good at agentic execution).

I also find the model is a lot more predictable and less “glitchy” when made to think in Chinese. You can do this in the system prompt.

anana_5 hours ago

Apparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219

Too little too late imo

akazantsev9 hours ago

> Gemma isn’t very good at agentic execution

I had no issues with it for C++ development with https://pi.dev. I'm yet to try it with Zed Editor. I don't rely on agents too much. However, I used it on Chromium's codebase to research some functionalities, let's say for searching. Requests like: check my last commit and do the same for SetterA and SetterB; it also ran without any errors.

wajahatraja11 hours ago

[dead]

rhdunn9 hours ago

Does anyone know if they intend on releasing open source/weights variants for 3.8 or whether 3.6 was the last model they are/were doing that for?

rolls-reus9 hours ago

will be releasing weights per their tweet announcing the model https://xcancel.com/Alibaba_Qwen/status/2078759124914098291

softwaredoug3 hours ago

It feels like an inflection point of lost US leadership in technology? A year plus ago you would say while China led in green energy and manufacturing, at least the US was ahead in software - as demonstrated by the state of US AI models.

We could point at a lot of factors on the US side. From political paralysis / head-in-the-sand attitudes towards emerging tech like green energy. To something of disdain for workers that will be impacted by AI (creating a backlash). To education that continues to lag. Add to this so many other self-inflicted economic wounds from the current administration.

I don't know if its nearly as terminal, as say the UK after WW2. The US is still large, wealthy, and resource rich. Yet at a minimum the triumphalism about US leadership after Trump was elected by the tech elite feels silly in retrospect.

Something I also think about is how much stronger The West overall would be if instead of antagonizing allies, there was a single ecosystem working closer together.

MichaelNolan7 hours ago

If 3.8 max goes open weight, what are the odds they retroactively open weight the earlier releases?

Gecko40727 hours ago

What would be the point? Kimi is kind of forcing their hand.

kennywinker7 hours ago

Well, if nothing else, posterity.

netdur10 hours ago

The only problem I had with Qwen, fine tuning on Colab, it takes 31 t/s while Gemma 4 is around 9 t/s, otherwise, one of best local LLM

brunooliv13 hours ago

Qwen is so much better than GLM or Kimi that this makes me genuinely excited!!

Pesto13 hours ago

The bigger models are usually worse though, hopefully they nail it this time.

try-working11 hours ago

in case someone wants to speculate in why chinese labs open source their models: https://try.works/why-chinese-ai-labs-went-open-and-will-rem...

sbinnee10 hours ago

If it offers more than opencode go, the entry plan looks enticing

sampton4 hours ago

This is reminiscent of the operating system wars and browser wars. In the end there will only be 2 models that can survive. 1: give it away for free or 2: locked in with top notch hardware.

nullbio11 hours ago

I predict that no one will use this and everyone will use Kimi K3.

embedding-shape11 hours ago

I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.

sunaookami10 hours ago

Same problem with every chinese model currently, they overthink way too much and take too much tokens and time.

embedding-shape10 hours ago

More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)

EgregiousCube10 hours ago

A consequence of aggressive distillation?

szundi10 hours ago

[dead]

rubslopes10 hours ago

Why? Price? If the reason is performance, I've been using non-frontier models for cheap, and they run great for my needs (GLM 5.2, DeepSeek v4 Pro).

rurban9 hours ago

We'll probably use it, but for images. Qwen is still the best for images

jadbox10 hours ago

What's the price difference?

corv9 hours ago

Who is behind this site? Is this another frontend to Alibaba or a reseller in Singapore?

raised_hand5 hours ago

interesting, when will this race end?

dartharva5 hours ago

I very much appreciate the existence of these free models, but in my experience Qwen has too high of a tendency to confidently give the wrong answer as compared to other frontier models.

eurekin12 hours ago

With 3.6 27b, I just stopped changing local models and started tinkering with things on top (like mem0). Feels genuinely useful and more than a toy

androiddrew10 hours ago

I have only been using 3.6 27B for coding. Is mem0 for agents like Openclaw or Hermes? How are you using it?

eurekin9 hours ago

It's a mcp, so connects quite easily to agents. With mcpo, I also connected it to open-webui (which has better support for OpenAI style tools/functions). Used it in claude code with that mcp plugin set-up too. Only ever used it for managing homelab information, but it met initial expectations. 27b is a great model, if grounded. The query about physical hosts and routing... I haven't found a single hallucination (altough Codex 5.6 as a reviewer mentioned something was wrong with some parts, and those were exactly the never properly documented ones. Codex/gpt had extra knowledge, because it was the conversation I used to set it up).

baist012 hours ago

can i get "code instruct" version of this? i want 7B and 14B to launch on my hardware.

ernsheong11 hours ago

So are locally-runnable models frozen at Qwen 3.6 now :/

worldsavior11 hours ago

Everyone wanted open models that would challenge Opus and Codex, here, you got it.

ernsheong10 hours ago

We need better coding models that can run on local hardware, i.e. 128GB VRAM or less

zozbot2349 hours ago

You can run larger models by offloading to SSD (for weights), it's just slow so people don't do it all that much. But you can get back at least some of that performance by using either MTP (at least for dense models; not effective for sparse MoE models unless you're batching them already and have VASTLY more parallel compute than you'd know what to do with) or batching multiple requests in parallel (note, this hurts throughput for your single sessions but running more sessions in parallel still boosts your total amount of inference. This requires careful management of memory requirements for your context/KV cache, and Qwen models tend to be KV-cache heavy).

Broadly speaking, this ultimately pushes local inference towards a challenging world where you use SSD offload for weights as a matter of course; then smaller requests (or requests sharing the bulk of their context, e.g. subagent swarms) can be batched together and run quickly in aggregate, but running very large contexts will actually limit you to single-session inference and require swapping out even the KV cache itself to some external scratch SSD, further hurting your performance. Then feel free to add wide use of MTP in a probably futile effort to go back to tolerable tok/s numbers.

seanmcdirmid10 hours ago

Queen has that already, although they seem to be moving away from local models unfortunately.

tormeh11 hours ago

Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.

ch_sm11 hours ago

In my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet.

mark_l_watson11 hours ago

I found qwen3.6:26b slightly better on my 32G mac mini than the same sized gemma until gemma was updated with better tool support 4 or 5 days ago.

It is like a ping-pong game: the advantage flips back and forth between providers.

regularfry10 hours ago

Worth knowing that Unsloth have just put out another Gemma 4 release from Google's upstream updates which should improve reliability. Bugs in the chat template affecting tool calling and other issues, apparently. https://www.reddit.com/r/unsloth/s/MpC6Hzs4Wj

dofm10 hours ago

Wow, thanks. I didn't see Unsloth had already done their version; I was just about to go back to the google version to test this change.

androiddrew10 hours ago

I have been running 3.6 27b on a dual AMD r9700 setup using Opencode and Matt Pocock's skills workflow for writing Golang CLIs. It's decent, but won't win any awards on code architecture. I guess you can try to AGENTS.md the deficits but I am just exploring its raw Opencode experience right now. Much slower than an API but still 3x times faster than I can read. Tuning it in with a community chat template and a specific penalty for repeats was the sauce needed to get it to work. I can probably start loop daddying it now over the tickets Matt's flow creates.

So yeah, it's the best local model I've seen. I am going to try the Qwopus 3.6 fine tune soon with the same spec and tickets and compare the output of both.

SomeHacker449 hours ago

Would you mind sharing more please? I literally just finished the same set up, with a 9950X CPU and 192G RAM at 4,800 MT/s. I used lemonade with Vulkan and the UD-Q8_x_x model from HF. 256k context. I have about 8G VRAM free, and use the iGPU for my desktop/monitor on Arch. What options do you give llama-cpp or whatever you run please? What other models have you found fit nicely in the 2xR9700? Thanks!!

seanmcdirmid10 hours ago

I actually have long discussions with Gemini about this and have wound up download a bunch of different models for different things. There is no best, just fast but worse, slow but better, agentic or not, reasoning or not great at large contexts, better world knowledge, uncensored, etc…. It’s a bit daunting actually since there isn’t really a one size fits all model that you can just use for everything.

ernsheong10 hours ago

Yes it's between this and Gemma 4 31B which is much slower, but looks like it won't ever get an upgrade. I have to conclude that the MoE variants are unreliable, and MTP sometimes just can't get tricky formatting right.

dofm10 hours ago

The whole series had an upgrade a couple of days ago actually — they have addressed embedded tool calling (and hopefully the MTP formatting stuff though I gave up running the Gemma MTP because it's often slower than not-MTP)

Not tried it yet but I've seen tests that suggest they've properly fixed the tool calling issues.

SwellJoe10 hours ago

I find the 4-bit QAT with MTP to be entirely usable speed on both my boxes (Strix Halo and a desktop with two V620 GPUs, which are slightly faster than the Strix Halo).

cmrdporcupine10 hours ago

For whatever reason prefill (on my DGX Spark) is faster with the Gemma models than Qwen 3.6 models of similar size. On vLLM anyways. Likely just deeply tuned code contributed to vLLM by Google?

vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE.

hnfong11 hours ago

People have been able to run DeepSeek v4 flash with a high spec Mac.

schaefer9 hours ago

I flip flop between qwen 3.6 27b and qwen 3.6 35b 4b active.

But there’s also the quantization of DeepSeek v4 flash called dwarfstar

cmrdporcupine11 hours ago

Gemma4 models are arguably better. Or at least about the same.

atemerev10 hours ago

The best model you can run locally is Kimi K3, as long as you have the hardware. If "what model I can still run on a something resembling something I can put on desktop without separate electricity and cooling water inputs", then it is probably GLM 5.2 (can be run on e.g. Nvidia DGX Station workstation). As long as you have about $100k-$150k.

dofm10 hours ago

Maybe, maybe not. Qwen 3.6 27B is literally just three months old. Hard to predict. Maybe it just wasn't worth making a 3.7, and after all, the 27B release was after the Plus release.

khurs13 hours ago

Go China, screw America*

*within the scope of open models only

dannyw13 hours ago

I like my Apache 2.0 licensed Gemma, and NVIDIA’s Nemotrons are decent bases for finetuning or continued pretraining, esp thanks to good documentation and tooling.

Oh, and Mira’s thinking machines lab dropped Inkling, a ~1T open weight model too.

This isn’t US vs China. This is open vs closed.

jimbob456 hours ago

This isn’t open though. Promises to be open later aren’t worth anything, given what we’ve seen and heard from AI execs making promises in this industry.

khurs12 hours ago

It's Sunday morning so I'm allowed to be facetious!

gxs12 hours ago

The open weights vs frontier models is reminding me more and more of the Linux vs Windows I grew up with (slashdot randomly popped into my head saying that)

I have a feeling this is the next…frontier of that fight

One can only hope it eventually does as well as Linux

Archit3ch9 hours ago

Obligatory "Does it answer security questions?".

mannanj5 hours ago

It's an interesting time to be alive when your local models are supposedly the pinnacle of what a free nation is capable of, yet the ethicality of the companies is disliked and their models restrict and limit you so much you root for the models from a socialist/communist state. If it wasn't for the effectiveness of propaganda, tribalism and psyops in this scenario my words wouldn't even be controversial and would just be seen as a truthful observation.

ludydev7 hours ago

[flagged]

souravsspace3 hours ago

[flagged]

bdxn9 hours ago

[dead]

jane_hilly9 hours ago

[dead]

jane_hilly13 hours ago

[dead]

hermes_scanner11 hours ago

[dead]

hunmernop9 hours ago

[dead]

adnane46 hours ago

[dead]

adnane46 hours ago

[dead]

adnane46 hours ago

[dead]

Umair_khan23247 hours ago

nice

dluan11 hours ago

waic go brr

cadlernox9 hours ago

Nice

nwhnwh13 hours ago

Open what?

smnplk10 hours ago

Do this giant open-weight models have less active params and could be run on consumer hardware or no ?

adrian_b9 hours ago

Any open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC.

The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful.

Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per second.

With open weights models that you host yourself, you are not constrained to use any single model, because that is the one for which you pay a subscription.

You can use many models, each for whatever it is more suitable. You can use frequently a small model with a high inference speed, but for some tasks you may actually save time with a better model, even if it is much slower.

In my opinion, even at 1 token per second a big model may be useful for some tasks.

apitman7 hours ago

I can only imagine what that does to the poor SSD

Alpha30315 hours ago

Paging in experts is mostly reading so on the first order effects it would be fine. These might be second order effects from e.g. write caches needing to be flushed more often (and maybe swapping other applications, if you use a swap file or partition) but it probably wouldn't be too much of an issue.

vitorgrs10 hours ago

SVG's pelican https://gist.github.com/vitordelucca/521c2d63c9b852c622e7648...

Made on the website, so not sure if on the API there's more thinking options...

esrauch10 hours ago

I feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?

rhdunn9 hours ago

A parachuting flamingo? An aardvark driving a bus? It should be easy to randomize the animal and the mode of transport (or vary it with animal playing a sport) to create images not in the training data.

onlyrealcuzzo6 hours ago

How about an animated SVG of a pelican doing the Macarena, profile view, spinning to face the camera on the last beats?

rvz5 hours ago

> I feel like the pelican test can't be relevant anymore;

It never was. The point of this "pelican test" was for performative reasons, or just for attention of the joke.

It is like trying to test whether if an adult elephant could actually climb up a tree and reporting that some elephants are slightly better at doing that than others while also reporting at the same time that they are all bad at tree climbing anyway.

This is an example of testing for the sake of testing. The "pelican test" tests for nothing.

joegibbs10 hours ago

What about an armadillo playing a piano? There are so many potential combinations It would say something if the pelican looked great but the armadillo looked terrible

LatencyKills10 hours ago

Agree. It was interesting/fun for a bit though.

cakbeslik10 hours ago

Using QWEN models since 2.5. I never used the chat properly but as an API I can say they're quite good, especially when you compare with OpenAI models. Cheaper and almost same level. I will try this now also.

comandillos13 hours ago

Just imagine Anthropic making Opus open-weights now for the sake of trolling everyone. Wouldn't surprise me at this point xD

hodgehog1113 hours ago

That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.

anon37383912 hours ago

I’ve seen no evidence that he believes in anything. He comes off as just another slimy would-be monopolist to me.

cyanydeez10 hours ago

I see no evidence that any ceo retains anything but the desire to capitalize on their marketplace of ideas for their own benefit. Like wolves inn sheep clothing, they'll put on any skin suit to convince people to keep giving them money and power.

And it has nothing to do with the individual, from what I can tell, 70% of the population placed in their position would become the same type of uberpath.

cindyllm10 hours ago

[dead]

embedding-shape13 hours ago

That OpenAI releases Sol as downloadable weights feels way more likely than Anthropic releasing even the tiniest of models for download.

ferrouswheel13 hours ago

Has Anthropic released anything at all, ever?

Kuxe2 hours ago

Not voluntarily

kingstnap4 hours ago

Anthropic has a lot of interpretability work, but they are extremely defensive about everything else. Dario doesn't really believe in releasing anything. Despite how he acts in terms of some sort of highly principle driven saviour (machines of loving grace) he clearly is more business minded then anything else.

For example they don't even tell you anything about the tokenization. They even do random chunking and padding to avoid leaking the token strings in the streaming api after it got reverse engineered. (See: https://spylab.ai/blog/claude-tokenizer/)

Alifatisk11 hours ago

Model Context Protocol

anonym2911 hours ago

safety datasets and a lot of safety related research