How to Set Up Fast Inference for Voice and Multimodal Models with SGLang-Omni
SGLang-Omni provides a specialized runtime for multi-stage inference of voice and multimodal models, handling the complex pipeline from audio encoding to speech synthesis.
Language
HomeLanguages
Sections
SGLang-Omni provides a specialized runtime for multi-stage inference of voice and multimodal models, handling the complex pipeline from audio encoding to speech synthesis.
Turn an old Android phone into a smart AI surveillance system with Sentinel. Captures video over LAN, records to PC, and uses multimodal neural networks for real-time event recognition.
Tired of juggling scattered training scripts for multimodal models? lmms-engine is a modular engine that handles distributed training, sequence packing, and GPU optimizations so you can focus on config and data.
A plugin for Claude Code that converts video into frames and timestamped transcripts, letting the AI analyze screen recordings and YouTube videos directly in the terminal.
Firebase team open-sourced Genkit, a framework for building AI features. Does it help combine Gemini, OpenAI, and Ollama without spaghetti code?
A practical guide to building AI agents that actually work. The ai-agent-book repository by Bojie Li covers everything from context engineering to self-evolution without fine-tuning.
MLX-Audio brings lightning-fast TTS, STT, and STS capabilities to Apple Silicon Macs. Built on Apple's MLX framework, it offers Whisper, Kokoro, voice cloning, and more—all running locally without cloud costs.
AnythingLLM transforms your documents into a smart knowledge base you can chat with. Supports 20+ LLMs, no-code agents, and multimodal processing.