All articles
Ten Tracks a Day: Kominami's AI Music Workflow

Production notes

Ten Tracks a Day: Kominami's AI Music Workflow

Kominami searches for unfamiliar genre combinations in Suno, rejects weak outputs, pairs selected tracks with matching or deliberately mismatched ChatGPT artwork, and hands the repetitive work after FFmpeg to software. This is the real workflow behind publishing ten tracks a day.

Searching for genres I have not heard before

The process begins by generating music in Suno. Rather than reproducing one established genre faithfully, I build the Styles text to push instruments, tempo, texture, era, and regional influences into combinations that feel unfamiliar or difficult to name.

Many outputs do not work. Some contain broken sounds, lose their appeal halfway through, or settle into something too familiar. I discard them and continue until I find tracks worth keeping. The central human task in this workflow is repeated direction and selection.

The rule of ten tracks a day

I have a personal rule to publish ten tracks to YouTube each day. I continue the search until the day's set is complete. The number is not only an output target; it forces me to reach combinations I would never find if I stopped after the first few acceptable results.

A single track could always receive more time. Kominami also treats the breadth of the catalog as part of the work. Instead of waiting indefinitely for one perfect track, I make repeated attempts in unknown directions and allow unexpected music to emerge from the volume of experiments.

Artwork that matches—or deliberately does not

For each selected track, I generate artwork with ChatGPT. Sometimes the image follows the track's mood. At other times I choose a deliberate mismatch, such as a bright image for dark music, because the tension between sound and image can invite a different interpretation.

Choosing images for ten tracks is tiring, especially after the music has already required extensive selection. I sometimes accept an image once it feels interesting enough rather than perfectly matched. That compromise is part of sustaining a daily practice.

After FFmpeg, the software takes over

I give the chosen MP3 and PNG the same name and place them in an incoming folder. The pipeline waits for both files to become stable, uses FFmpeg to render a still-image video, stores the audio and artwork in Cloudflare R2, and uploads the video to YouTube as private before scheduling it.

It then matches titles and filenames, registers the track in the database, analyzes several sections of the audio with local classifiers, combines those results with the saved Suno Styles and catalog lane, and assigns candidate genres, moods, and use cases. It also prepares the YouTube description and download page, then separates successful files from failures.

  • Pair identically named MP3 and PNG files
  • Render the still-image video with FFmpeg
  • Store audio and artwork in R2
  • Upload privately to YouTube and schedule publication
  • Register the track and add analyzed metadata
  • Separate successful assets from failures

Automating what comes after creation

The difficult part is not pressing Generate. It is deciding what unfamiliar direction to try, listening, rejecting weak material, selecting ten tracks, and deciding whether each image should match or resist the music.

If rendering, uploading, sorting files, and database registration also remained manual, there would be less energy for the next round of judgment. The software therefore automates the logistics after creation, not the creative decisions that define what enters the catalog.

Continuity over perfection

This workflow combines generation, selection, automation, and compromise. Not every track will be received equally, and I may later question some choices. I still continue. Kominami does not decide the answer in advance; it creates, selects, publishes, and looks for discoveries inside the accumulated results.