Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
M Veselovskiy
Yuuru
28
2
11
Follow
21world's profile picture
1 follower
·
9 following
AI & ML interests
None yet
Recent Activity
liked
a model
3 days ago
Qwen/Qwen3.8-27B
new
activity
over 1 year ago
mistralai/Mistral-Small-3.1-24B-Instruct-2503:
HF Format?
reacted
to
m-ric
's
post
with 👀
over 1 year ago
𝗔𝗱𝘆𝗲𝗻'𝘀 𝗻𝗲𝘄 𝗗𝗮𝘁𝗮 𝗔𝗴𝗲𝗻𝘁𝘀 𝗕𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸 𝘀𝗵𝗼𝘄𝘀 𝘁𝗵𝗮𝘁 𝗗𝗲𝗲𝗽𝗦𝗲𝗲𝗸-𝗥𝟭 𝘀𝘁𝗿𝘂𝗴𝗴𝗹𝗲𝘀 𝗼𝗻 𝗱𝗮𝘁𝗮 𝘀𝗰𝗶𝗲𝗻𝗰𝗲 𝘁𝗮𝘀𝗸𝘀! ❌ ➡️ How well do reasoning models perform on agentic tasks? Until now, all indicators seemed to show that they worked really well. On our recent reproduction of Deep Search, OpenAI's o1 was by far the best model to power an agentic system. So when our partner Adyen built a huge benchmark of 450 data science tasks, and built data agents with smolagents to test different models, I expected reasoning models like o1 or DeepSeek-R1 to destroy the tasks at hand. 👎 But they really missed the mark. DeepSeek-R1 only got 1 or 2 out of 10 questions correct. Similarly, o1 was only at ~13% correct answers. 🧐 These results really surprised us. We thoroughly checked them, we even thought our APIs for DeepSeek were broken and colleagues Leandro Anton helped me start custom instances of R1 on our own H100s to make sure it worked well. But there seemed to be no mistake. Reasoning LLMs actually did not seem that smart. Often, these models made basic mistakes, like forgetting the content of a folder that they had just explored, misspelling file names, or hallucinating data. Even though they do great at exploring webpages through several steps, the same level of multi-step planning seemed much harder to achieve when reasoning over files and data. It seems like there's still lots of work to do in the Agents x Data space. Congrats to Adyen for this great benchmark, looking forward to see people proposing better agents! 🚀 Read more in the blog post 👉 https://huggingface.co/blog/dabstep
View all activity
Organizations
Yuuru
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
mistralai/Mistral-Small-3.1-24B-Instruct-2503
over 1 year ago
HF Format?
🧠
❤️
33
41
#2 opened over 1 year ago by
bartowski
New activity in
TheDrummer/UnslopSmall-22B-v1-GGUF
almost 2 years ago
Metharme format makes model extremely stupid
2
#1 opened almost 2 years ago by
Ainonake
New activity in
mattshumer/Reflection-Llama-3.1-70B
almost 2 years ago
DLETE THIS MODEL
👍
9
2
#76 opened almost 2 years ago by
MaziyarPanahi
New activity in
yodayo-ai/kivotos-xl-2.0
about 2 years ago
Broken results
3
#1 opened about 2 years ago by
Yuuru
New activity in
TheBloke/dolphin-2.6-mistral-7B-GPTQ
over 2 years ago
I am having issue loading dolphin 2.6 mistral 7B GPTQ:main by TheBloke . Pasting the error in the description pls help
❤️
2
3
#1 opened over 2 years ago by
KingSlayer49
New activity in
chargoddard/mixtralnt-4x7b-test
over 2 years ago
It works!!!
❤️
1
7
#1 opened over 2 years ago by
HoangHa
New activity in
TheBloke/Mixtral-8x7B-v0.1-GGUF
over 2 years ago
It works.
👍
🤗
2
6
#3 opened over 2 years ago by
Yuuru
New activity in
mistralai/Mistral-7B-Instruct-v0.2
over 2 years ago
How is this different from v1?
👍
9
7
#2 opened over 2 years ago by
amgadhasan
New activity in
stabilityai/stable-video-diffusion-img2vid-xt
over 2 years ago
Did anyone figured it out how to run it in low vRAM like 15-25 GB
5
#21 opened over 2 years ago by
zohadev
New activity in
TheBloke/WizardLM-30B-Uncensored-GGML
about 3 years ago
WizardLM-30B-Uncensored - reasoning is on the level of 65B models !
❤️
👍
5
3
#2 opened about 3 years ago by
mirek190
New activity in
TheBloke/falcon-40b-instruct-GPTQ
about 3 years ago
GGML?
7
#2 opened about 3 years ago by
creative420
New activity in
TheBloke/WizardLM-Uncensored-Falcon-7B-GPTQ
about 3 years ago
Found a script to convert to GGML
7
#2 opened about 3 years ago by
algorithm
New activity in
Aitrepreneur/stable-vicuna-13B-GPTQ-4bit-128g
over 3 years ago
Completely hallucinating
2
#1 opened over 3 years ago by
Rbonnav
New activity in
eachadea/ggml-vicuna-13b-1.1
over 3 years ago
ggml-vicuna-13b-1.1-q4_3 unrecognized tensor type 5
6
#10 opened over 3 years ago by
NicRaf
New activity in
Neko-Institute-of-Science/LLaMA-7B-4bit-128g
over 3 years ago
Does it req Triton to use?
2
#1 opened over 3 years ago by
Yuuru
New activity in
anon8231489123/vicuna-13b-GPTQ-4bit-128g
over 3 years ago
Vram usage
11
#3 opened over 3 years ago by
Juuuuu
New activity in
AlekseyKorshuk/vicuna-7b
over 3 years ago
Only 13b has been released, how can this be 7b?
4
#1 opened over 3 years ago by
underlines