With this fu di g, Modulate pla s to i c ease i vestme t ac oss AI/ML esea ch, p oduct a d e gi ee i g, develope elatio s, a d pa t e ships, while co ti ui g to b oade the APIs, models, a d deployme t optio s available to develope s.
The i vestme t follows a pe iod of sig ifica t tech ical a d comme cial mome tum fo Modulate. Its models ow a alyze mo e tha 10 millio hou s of audio each mo th, ece tly su passi g ove 600 millio hou s of audio p ocessed i total, while its t a sc iptio a d deepfake detectio tech ologies both a ked #1 o public be chma ks like Huggi g Face, the i dust y’s autho itative platfo m fo leade boa ds, evaluatio datasets, a d sta da dized model testi g.
Modulate’s audio- ative models a e used eve y day to p otect healthca e i stitutio s f om deepfake hacke s; e ha ce voice AI age t’s emotio a d empathy capabilities; educe ext emism a d ha assme t o social platfo ms; obse ve a d mo ito voice age t pe fo ma ce; detect a d stop child g oomi g voice co ve satio s; a d p otect age ts i high- isk sce a ios f om bei g ide tified th ough adva ced voice maski g – amo g a g owi g umbe of use cases. Modulate is ow expa di g the team a d i f ast uctu e eeded to meet g owi g dema d f om develope s a d pa t e s buildi g voice applicatio s ac oss secu ity, custome expe ie ce, commu icatio s, AI age t supe visio , a d t ust a d safety.
“Voice is becomi g a p ima y i te face fo AI, a d that c eates a whole ew set of p oblems that ca ‘t be solved f om a t a sc ipt,” said Ca te Huffma , CEO a d co-fou de of Modulate. “We’ e al eady usi g audio- ative AI to p otect o ga izatio s f om deepfake attacks, help voice age ts u de sta d emotio a d espo d with mo e empathy, ide tify da ge ous behavio i o li e co ve satio s, a d mo ito whethe voice age ts a e actually pe fo mi g the way they’ e supposed to. U de eath all of that a e mo e tha a hu d ed specialized models wo ki g togethe to u de sta d what’s eally happe i g ac oss audio, with d amatically less cost a d compute tha t aditio al la ge models.”
As i vestme t pou s i to AI systems that ca speak atu ally, Modulate is focused o the othe side of the i te actio : helpi g machi es accu ately u de sta d what is happe i g i a voice co ve satio .
Modulate’s flagship Velma platfo m is the leadi g model fo u de sta di g co ve satio s, with 2x g eate accu acy tha t aditio al LLMs at detecti g t ue positive esults a d 7x fewe false positive esults. Velma a alyzes audio di ectly to ide tify sig als i cludi g emotio , to e, i te t, emphasis, sy thetic speech, a d co ve satio al behavio s. Those sig als ca be used i depe de tly o composed to ecog ize highe -level eve ts, f om f aud attempts a d AI age t failu es to ha assme t, custome dissatisfactio , a d policy violatio s. Velma ca ope ate i eal time, e abli g applicatio s ot o ly to u de sta d what happe ed i a co ve satio , but to i te ve e while it is still happe i g.
The u de lyi g tech ology powe i g Velma is Modulate’s E semble Liste i g Model a chitectu e, o ELM. Rathe tha elyi g o a si gle massive fou datio model, Modulate’s ELM o chest ates mo e tha 100 specialized audio models, selecti g a d combi i g them to delive highly accu ate esults while substa tially educi g the compute equi ed fo i fe e ce. Velma has demo st ated up to 1,000x g eate efficie cy tha a si gle la ge model app oach, educi g the cost, e e gy a d memo y equi ed to a alyze audio at scale.
With mo e tha 600 millio hou s of audio a alyzed, that tech ical app oach is al eady p oduci g measu able esults. Modulate ece tly ea ed the umbe -o e positio o Huggi g Face’s Ope ASR Leade boa d fo t a sc iptio a d cu e tly a ks fi st o Huggi g Face’s deepfake speech be chma k. Its t a sc iptio API is p iced at $0.03 pe hou fo batch p ocessi g, while its deepfake detectio tech ology achieves 98.9% accu acy o public be chma k data.
Modulate’s voice- ative tech ology is al eady solvi g a expa di g set of eal-wo ld p oblems. Its models ca help o ga izatio s detect sy thetic voices a d suspicious behavio i high- isk calls; ide tify whe a voice AI age t is f ust ati g o misu de sta di g a custome ; detect ha assme t, g oomi g a d othe ha mful behavio o social a d gami g platfo ms; a d give develope s audio- ative sig als that allow AI applicatio s to espo d mo e app op iately to the people usi g them.
Steve Ju vetso , Co-fou de of Futu e Ve tu es a d Boa d membe of SpaceXAI, said, “Modulate has gai ed a sig ifica t tech ical lead i audio- ative AI, a d the ma ket oppo tu ity is expa di g quickly. The team has p ove these models i some of the most dema di g voice e vi o me ts i the wo ld, a d we’ e ow seei g the eed fo that tech ology to expa d well beyo d whe e it sta ted i to AI age ts, secu ity, custome expe ie ce, a d mo e. This i vestme t will help Modulate move faste , g ow the team, put its models i to the ha ds of mo e develope s a d pa t e s, a d establish audio i tellige ce as a fou datio al laye of the AI stack.”
The fu di g will also suppo t Modulate’s g owi g develope a d pa t e st ategy. The compa y is buildi g ew i dust y models, expa di g its develope tooli g with ew SDKs a d APIs, i c easi g develope elatio s esou ces, c eati g pa t e i teg atio s, a d suppo ti g ew deployme t e vi o me ts fo custome s. With these additio al esou ces, Modulate is givi g compa ies buildi g voice age ts, commu icatio s platfo ms, secu ity p oducts, a d othe audio applicatio s a way to i co po ate sophisticated audio u de sta di g without developi g specialized models themselves.
“Develope s should ‘t have to ebuild the audio i tellige ce laye eve y time they c eate a ew voice expe ie ce,” Huffma added. “Ou missio is to build the models a d i f ast uctu e that let them focus o the applicatio they wa t to c eate. The oppo tu ity faci g audio- ative AI is expa di g i c edibly quickly. We’ve built the tech ology a d p ove it at scale, a d this i vestme t lets us g ow the team a d move faste to meet that dema d.”
With the ew capital, Modulate is scali g the team, tech ology a d develope ecosystem behi d its b oade ambitio : maki g audio- ative i tellige ce a fou datio al laye whe eve voice AI is built.
Dow load the Modulate p ess kit he e.
Modulate is a f o tie audio AI compa y buildi g audio- ative models that powe the ext ge e atio of voice AI, e abli g machi es to u de sta d the ua ces i huma co ve satio beyo d the wo ds bei g spoke . Its Velma platfo m is powe ed by Modulate’s E semble Liste i g Model (ELM) a chitectu e, b i gi g togethe mo e tha 100 specialized models to u de sta d sig als i cludi g emotio , to e, i te t, sy thetic speech a d co ve satio al behavio a d c eate a powe ful laye of audio i tellige ce fo voice applicatio s. Fou ded by MIT alum i, Modulate’s tech ology has a alyzed mo e tha 600 millio hou s of audio a d is used ac oss AI age ts, f aud a d deepfake detectio , custome expe ie ce, t ust a d safety, a d othe eme gi g voice applicatio s.
Fo mo e i fo matio o to get sta ted, visit modulate.ai.
K isti Ca de sG ithaus Age cy207-974-7744(e)
###






 