“O e of ou ea ly custome s, a leadi g global tech ology compa y, told us we’ve eached a si gula ity mome t fo ei fo ceme t lea i g with huma feedback, whe e evide ce-g ou ded automatio c osses the huma -quality th eshold at scale,” said Law e ce S app, CEO of T ustScale. “This fu dame tally cha ges the eco omics of AI t ai i g a d ei fo ceme t lea i g. AI make s a d deploye s o lo ge have to choose betwee the scale of automatio a d the quality of huma evaluatio . They ca have both, g ou ded i dete mi istic evide ce athe tha a othe p obabilistic AI opi io .”
Huma feedback has lo g bee the gold sta da d fo evaluati g a d imp ovi g AI models th ough ei fo ceme t lea i g. As AI developme t accele ates, model make s a e i c easi gly automati g that p ocess with AI judges a d othe model-based evaluatio systems. A gusRL takes a diffe e t app oach, usi g empi ical evide ce a d dete mi istic ve ificatio to evaluate AI-ge e ated espo ses a d ge e ate st uctu ed feedback that ca be used to co ti uously imp ove model pe fo ma ce.
I a p oductio deployme t with a leadi g global tech ology compa y, A gusRL’s automated evaluatio delive ed bette esults tha the custome ‘s huma a otato s a d ide tified mistakes the huma eviewe s had missed.
“We’ve wo ked with a a ge of pa t e s a d app oaches to imp ove the quality of ei fo ceme t lea i g a d model evaluatio , a d A gusRL has co siste tly stood out fo the quality a d accu acy of its p ompt a d espo se eview,” Fo me Apple a d Amazo AGI Leade
U like AI-Judge app oaches that ely solely o p obabilistic AI to evaluate a othe p obabilistic system, T ustScale’s pate t-pe di g A gusRL tech ology g ou ds its evaluatio s i et ieved exte al evide ce. The platfo m a alyzes each p ompt a d espo se, b eaks espo ses i to i dividual claims a d sea ches multiple data sou ces fo suppo ti g o co t adicto y evide ce. It the etu s st uctu ed dete mi istic ve dicts with citatio s a d co fide ce sco es.
A gusRL also evaluates the quality of the o igi al que y a d ove all espo se a d ide tifies cases that wa a t huma eview. With A gusRL, co t adicted claims, claims without sufficie t evide ce a d othe flagged espo ses ca be outed to huma a otato s, allowi g people to focus o the cases whe e huma judgme t adds the g eatest value athe tha ma ually evaluati g eve y espo se.
Because A gusRL co ti uously evaluates outputs afte deployme t, its ei fo ceme t feedback ca i co po ate cu e t evide ce a d i fo matio that may ot have bee available du i g a model’s i itial t ai i g.
“The implicatio s go well beyo d accu acy,” said S app. “AI compa ies spe d billio s of dolla s each yea o the data, huma evaluatio a d i f ast uctu e equi ed to t ai a d imp ove models. If AI make s ca automate mo e of the ei fo ceme t feedback p ocess without sac ifici g quality, they ca imp ove models faste a d at lowe cost while ese vi g huma expe tise fo the cases that actually equi e it.”
A gusRL is built o the same evide ce-based T ustScale E gi e that powe s A gus, the compa y’s AI assu a ce platfo m fo detecti g a d co ecti g halluci atio s at the poi t of use. A gusRL takes that evide ce-based app oach upst eam, givi g AI make s a d develope s st uctu ed feedback they ca i co po ate i to model t ai i g, fi e-tu i g pipeli e, a d co ti uous imp oveme t.
A gusRL ope ates as a API-backed evaluatio se vice a d suppo ts multiple la guages, locales a d i put fo mats. It ca be i teg ated with existi g model developme t, evaluatio , a d a otatio wo kflows a d etu s claim-level ve dicts, suppo ti g evide ce, citatio s, a d st uctu ed esults fo dow st eam use.
A gusRL is available today th ough the AWS Ma ketplace a d di ectly th ough T ustScale. To lea mo e o equest a demo st atio , visit https://t ustscale.ai/e /a gus l.
T ustScale is a AI t ai i g, evaluatio a d assu a ce compa y helpi g o ga izatio s c eate, shape a d use a tificial i tellige ce with g eate co fide ce a d co t ol. Built o mo e tha 20 yea s of expe ie ce i AI data ac oss 200-plus la guages, T ustScale develops i depe de t tech ologies that detect AI mistakes, evaluate claims agai st dete mi istic empi ical evide ce a d keep people at the ce te of co seque tial decisio s. Its A gus suite spa s the AI lifecycle, f om eal-time halluci atio detectio a d evide ce-based co ectio at the poi t of use, to automated evaluatio a d ei fo ced feedback fo model t ai i g a d co ti uous imp oveme t. Lea mo e at T ustScale.ai.
So gue PR fo T ustScale








 