Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Alex Kolo's picture

Alex Kolo

alexkolo
2 11
Β·
  • alexkolo
  • alexkolodzig

AI & ML interests

None yet

Organizations

Hugging Face Discord Community's profile picture

Collections 4

Parser
  • microsoft/OmniParser

    Image-Text-to-Text β€’ Updated Dec 2, 2024 β€’ 605 β€’ 1.71k
transcribe
  • pyannote/speaker-diarization-3.1

    Automatic Speech Recognition β€’ Updated May 10, 2024 β€’ 8.24M β€’ 3.74k
  • pyannote/segmentation-3.0

    Voice Activity Detection β€’ Updated May 10, 2024 β€’ 5.81M β€’ 1.88k
  • openai/whisper-large-v3

    Automatic Speech Recognition β€’ 2B β€’ Updated Aug 12, 2024 β€’ 4.76M β€’ β€’ 6.32k
Parser
  • microsoft/OmniParser

    Image-Text-to-Text β€’ Updated Dec 2, 2024 β€’ 605 β€’ 1.71k
transcribe
  • pyannote/speaker-diarization-3.1

    Automatic Speech Recognition β€’ Updated May 10, 2024 β€’ 8.24M β€’ 3.74k
  • pyannote/segmentation-3.0

    Voice Activity Detection β€’ Updated May 10, 2024 β€’ 5.81M β€’ 1.88k
  • openai/whisper-large-v3

    Automatic Speech Recognition β€’ 2B β€’ Updated Aug 12, 2024 β€’ 4.76M β€’ β€’ 6.32k
View 4 collections

spaces 1

pinned
Sleeping
2

Audio Transcript with Gemini 2.5

πŸ“Š

Using `gemini-2.5-pro-exp-03-25` to transcribe audio files

Jun 6, 2025

models 0

None public yet

datasets 0

None public yet
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs