이 영상에는 한국어 자막이 포함되어 있습니다. (자막 버튼을 켜시면 영어 자막도 선택하실 수 있습니다.)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ ORIGINAL STREAM — all credit to Sage Elliott
Watch the original: https://www.youtube.com/watch?v=FVHKbJOS-Qw
Channel: https://www.youtube.com/@sagecodes/
LinkedIn: https://www.linkedin.com/in/sageelliott/
X: https://x.com/sagecodes
This is a subtitled version shared with attribution. Please support the original creator.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ WHY I SUBTITLED THIS
I built an automated subtitling pipeline — transcription, proper-noun correction,
translation, and burned-in captions — and needed a hard test case for it.
A 65-minute unscripted livestream is exactly that: overlapping speech, long music
playback with no dialogue, live chat interruptions, and dozens of product names that
speech recognition routinely gets wrong.
Some of what the pipeline had to fix:
“flight” → Flyte (the orchestration platform)
“ASTEP” → ACE-Step
“Grady app” → Gradio app
“Texas Beach” → text-to-speech
“Quinn” → Qwen
“mutual co” → MuJoCo
“Jan LeCun” → Yann LeCun
One case had to be left alone: near the end, “getting a flight” really does mean an
airplane. A blind find-and-replace would have broken that line. Context still needs
a human.
Song lyrics are marked with ♪ so you can tell singing from speech.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ WHAT’S IN THE STREAM
Sage compares open-source text-to-music models side by side, generates songs live
from viewer requests, and is refreshingly honest about what still doesn’t work.
Models covered:
• ACE-Step 1.5 — the main focus (turbo vs large backend)
• MiniMax Music — released the night before the stream, heard live for the first time
• DiffRhythm
• Stable Audio — instrumental only
• MusicGen (large)
• AudioLDM 2 — older, noticeably weaker
• Suno — commercial reference point
The most useful finding: duration matters more than almost any other parameter.
Cram lyrics into a short song and the vocals sound compressed and rushed. Give the
model room and it not only articulates better — it adds ad-libs that were never in
the lyrics.
Also discussed: what “sounding too AI” actually means (compressed vocal range,
especially on lower notes and harmonies), why drums and synths land but guitars
often don’t, licensing before commercial use, and the ACE-Step training set
(5M curated audio samples, Gemini 2.5 for lyric extraction).
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ CHAPTERS
00:00 Intro — wrapping up the generative media series
00:58 What’s next: RL, MuJoCo, world models
02:14 Setup: Flyte, DGX Spark, Hugging Face
03:48 Today’s topic: open-source music generation
04:08 ACE-Step vs Suno
04:58 ACE-Step 1.5 — first synthwave sample
06:57 The model lineup
07:28 Parameters: diffusion steps and guidance
08:24 The biggest factor is duration
09:12 Clip-by-clip model comparison
11:58 Check the license before you ship
12:41 Genre comparison
15:47 ACE-Step vs MiniMax — first listen
18:05 Outlaw country
19:57 Norwegian black metal
21:52 What does “sounding too AI” mean?
24:24 Training data — the ACE-Step paper
25:56 Negative prompts and audio-to-audio
27:45 Generating live with the chat
29:56 Electronic / Odesza-style
34:05 Recap: image, video, and TTS models
37:41 Cinematic heavy metal
40:11 Indie pop, sea shanty, and more
43:24 Mid-stream verdict
44:50 1980s arena power ballad
49:43 Lyrics vs song length experiments
52:41 Closing thoughts on the models
55:13 Next up: MuJoCo and reinforcement learning
58:32 Seattle Hack Night — RAG and vector stores
63:18 Wrap-up
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ RESOURCES (from the original stream)
GitHub: https://github.com/sagecodes/ai-build-and-learn
Events: https://luma.com/ai-builders-and-learners
Slack: https://slack.flyte.org/
ACE-Step 1.5: https://huggingface.co/ACE-Step/Ace-Step1.5
ACE-Step paper: https://arxiv.org/abs/2602.00744
Flyte: https://github.com/flyteorg/flyte
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
▶ ABOUT THIS CHANNEL
Catch Up AI — practical experiments in using AI for everyday work.
I’m a retired IT engineer (15 years in Korea, 15 in the US) figuring out how to
actually put these tools to work, and sharing what breaks along the way.
#AIMusic #ACEStep #OpenSourceAI #MiniMax #TextToMusic #AISubtitles
| Play | Cover | Release Label |
Track Title Track Authors |
|---|