Thanks to visit codestin.com
Credit goes to github.com

Skip to content
View xzf-thu's full-sized avatar

Block or report xzf-thu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. gpt-omni/mini-omni gpt-omni/mini-omni Public

    open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.

    Python 3.6k 310

  2. gpt-omni/mini-omni2 gpt-omni/mini-omni2 Public

    Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities。

    Python 1.9k 210

  3. Mega-ASR Mega-ASR Public

    First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come …

    Python 1.1k 75

  4. Pask Pask Public

    Towards Self-Evolving Proactive AI with Perpetual Memory

    Python 217 22

  5. Audio-Interaction Audio-Interaction Public

    Python 587 33

  6. VoiceMem VoiceMem Public

    Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing information and emotions respectively, while…

    Python 1.5k 108