this post was submitted on 01 Dec 2024
32 points (82.0% liked)

Technology

35117 readers
25 users here now

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.


Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.


Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

founded 5 years ago
MODERATORS
 

cross-posted from: https://futurology.today/post/2910566

Alibaba's Qwen team just released QwQ-32B-Preview, a powerful new open-source AI reasoning model that can reason step-by-step through challenging problems and directly competes with OpenAI's o1 series across benchmarks.

The details:

QwQ features a 32K context window, outperforming o1-mini and competing with o1-preview on key math and reasoning benchmarks.

The model was tested across several of the most challenging math and programming benchmarks, showing major advances in deep reasoning.

QwQ demonstrates ‘deep introspection,’ talking through problems step-by-step and questioning and examining its own answers to reason to a solution.

The Qwen team noted several issues in the Preview model, including getting stuck in reasoning loops, struggling with common sense, and language mixing.

Why it matters: Between QwQ and DeepSeek, open-source reasoning models are here — and Chinese firms are absolutely cooking with new models that nearly match the current top closed leaders. Has OpenAI’s moat dried up, or does the AI leader have something special up its sleeve before the end of the year?

top 8 comments
sorted by: hot top controversial new old
[–] [email protected] 32 points 2 weeks ago (3 children)

Let's give it a whirl!

welp

[–] [email protected] 2 points 2 weeks ago (1 children)

It works in Spanish, in English it throws an error before answering about Tiananmen as you show 🤫

[–] [email protected] 1 points 2 weeks ago

Interesting. I tried Chinese and it also throws an error. Looks like it was a manual thing in only some languages.

[–] [email protected] 1 points 2 weeks ago

Someone gagged the AI before it could complete that sentence 😜

[–] [email protected] 1 points 2 weeks ago (1 children)

I wonder if that's a UI block like if it's mentioned then throw error, or if the model itself has a block in there. From this, it looks like it's baked in and of course they haven't poisoned it

[–] [email protected] 1 points 2 weeks ago* (last edited 2 weeks ago) (1 children)

I'm guessing it's in the output handler, not the UI exactly. I don't think you can edit models like that, and the fact that it knows about it at all means they didn't whitewash the training data set. But my knowledge is limited. In their place, I would probably have included "don't talk about tiananmen square" in the initialization rules. But failing that, I would have added something in the output processor to check for forbidden knowledge and throw an exception.

Still, it's strange that it got the words out before dying.

[–] [email protected] 1 points 2 weeks ago

Yeah agreed, I'm more surprised they didn't scrub every reference to it on the training set like you said that it's in the model at all is surprising. I may try to run it myself and see what it does with the same question

[–] [email protected] -1 points 2 weeks ago* (last edited 2 weeks ago)

These data processing apps are just apps hooked up to big computers with big data. It's no surprise when rich people buy a bunch of computers and data then run an app on it. It's much more surprising that people are hyped into believing this is somehow important. "It took one week to copy an app and load the data?!?! Wow!!!!"

Nobody cares when "China" writes a decent word processing app or whatever, nor should they.