Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Tony He's picture

Tony He

ttttonyhe
4 10 26
Ruben148's profile picture tfrere's profile picture 6b4b86ec-928a-4b7e-9c1e-8d5f009e3272's profile picture
·
https://lipeng.ac
  • tonyhe_lipeng
  • ttttonyhe

AI & ML interests

Trustworthy Machine Learning

Recent Activity

updated a model 12 days ago
ttttonyhe/Qwen3-4B-Instruct-RETA
upvoted a paper 17 days ago
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
updated a collection 18 days ago
Prompt Injection
View all activity

Organizations

University of Waterloo's profile picture

New activity in huihui-ai/Huihui-Qwen3-4B-Instruct-2507-abliterated 7 months ago

fix: empty think blocks

#3 opened 7 months ago by
ttttonyhe
commented a paper 12 months ago

Locket: Robust Feature-Locking Technique for Language Models

Paper • 2510.12117 • Published Oct 14, 2025 • 2 •
2
commented a paper over 1 year ago

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Paper • 2502.00840 • Published Feb 2, 2025 •
3
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs