Voice cloning is an area where data ethics are paramount, and Metamorph set a high standard for how it should be done. I recruited and worked directly with the twelve singers behind its voice models, structuring each partnership to be fair, transparent, and mutually beneficial. I sourced the talent and personnel, coordinated the technical setup, and directed the recording sessions, supervising the signal chain to produce a consistent, artifact-free dataset of feature-rich performances, each with its own distinct character. Hear the results in Metamorph.
What does it take to teach a model to separate the human voice from everything else? Seven days, 100 artists, two recording studios, 54 microphones, 16 preamps, countless field recordings of birds, iPhones, and HVAC systems, and some seriously quiet signal chains. I took this project from idea to delivery and directed every session myself, showing that hands-on data collection can still work at scale. The result is the cleanest vocal de-noiser around.
A large phonetic annotation project where I served as the senior quality reviewer. Training a de-esser means teaching a model exactly where sibilance sits in a performance, which depends on precise, consistent phoneme labels across a big singing corpus. Working from an initial annotation pass, I applied a linguistically-minded QA layer that brought the labels to training quality, reviewing roughly 3,700 tasks and close to 13,000 individual phoneme annotations against the schema. My background in linguistics is what let me adjudicate the hard cases consistently, the affricates, the sibilant-plosive clusters, the ambiguous edges, so the training data was clean enough to trust. The result is a de-esser that is sophisticated under the hood and remarkably simple to use.