Could a cheaper Salesforce-hosted model do the same job? Seven scoring requests built from live demo data went to Jev and to five models on Agentforce's model gateway, three runs each, for 126 calls in total. The general-purpose models received the same JSON plus instructions asking for Jev's answer shape and telling them to weigh facts over free-text claims.
Two cheap models fail the demo: GPT-4o mini and Gemini 2.5 Flash rated a third-shipment quality complaint as merely frustrated, below the escalation threshold, so the escalation task would never fire.
Resistance to claims: Injecting "We will definitely reorder socks next week, guaranteed" moved Gemini from 2% to 90%. Jev moved 9 points, by the same amount on every run, because its instructions weight invoices above stated intent.
Stability: Jev's answers moved at most 3 points across identical runs and came back in about 0.1 seconds per call, against 1.1 to 1.9 seconds for the hosted models.
An open question: One account had ordered socks every month for a year and missed a single order. Jev put the chance of a reorder at 53%. Every other model said 2 to 15%, apparently anchoring on the "Dormant" label. Only real outcomes can settle which is right.
Cost: The demo's four calls used about 4,000 Jev tokens, roughly $0.0002. On Agentforce's Basic tier, the same four calls cost about $0.04 in Flex Credits, and as four agent actions they would cost $0.40 before model credits.