Google’s Gemini 3.6 Flash targets enterprise agent token prices


Google has launched Gemini 3.6 Flash and three.5 Flash-Lite as new workhorses designed to lower latency and token prices for enterprise AI brokers.

The economics of working autonomous software program brokers inside a manufacturing surroundings come down to a hard and fast equation few distributors promote instantly. A mannequin wants to cause by way of a multi-step process competently, however each additional token it generates whereas doing so provides price and delay to a workflow which may run hundreds of instances an hour.

Groups constructing background brokers somewhat than chat interfaces want throughput first and parameter rely second. Google’s reply, introduced this week, splits that trade-off throughout three fashions: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant constructed for vulnerability remediation.

The mathematics behind Gemini 3.6 Flash

Google’s developer documentation for 3.6 Flash centres on one determine: 17 % fewer output tokens than the prior 3.5 Flash model, based mostly on measurements from the Synthetic Evaluation Index.

In particular artificial checks, together with the Datacurve DeepSWE benchmark, Google studies drops in token utilization of up to 65 %. Pricing sits at $1.50/1M enter tokens and $7.50/1M output tokens, positioning the mannequin for reasoning loops that run repeatedly somewhat than on-demand.

On DeepSWE, the firm information a 49 % success price for 3.6 Flash towards 37 % for its predecessor. On MLE Bench, the rating strikes from 49.7 % to 63.9 %, and on Google’s GDPval-AA v2 take a look at – which makes an attempt to measure real-world data work somewhat than coding puzzles – 3.6 Flash scores 1421 towards 1349 for the older mannequin.

Figma, Hebbia, and Harvey put the mannequin to work

Figma has built-in 3.6 Flash into its prototyping infrastructure, and in accordance to Matt Colyer, the firm’s Director of Product Engineering, the mannequin provides builders a sooner route by way of design iterations with no drop in output high quality.

Authorized know-how platform Harvey and analysis instrument Hebbia route information by way of the mannequin for multimodal doc work: ingesting uncooked monetary filings, parsing doc construction, studying embedded charts, and producing draft studies for evaluation.

Google additionally folded a client-side computer-use instrument instantly into the Gemini API and Gemini Enterprise platforms, eradicating the customized middleman software program engineers beforehand constructed to let fashions function on prime of an working system.

The corporate studies an OSWorld-Verified rating of 83.0 %, up from 78.4 %, and says up to date safeguards towards chemical, organic, radiological, and nuclear misuse enhance resistance to jailbreaking with out elevating refusal charges for benign requests.

A less expensive tier for high-volume background brokers

Gemini 3.5 Flash-Lite targets a unique job: doc processing and agentic search working at quantity somewhat than reasoning depth. The Synthetic Evaluation Index measured the mannequin at 350 output tokens per second, the quickest in the 3.5 collection in accordance to Google.

Pricing runs at $0.3/1M enter tokens and $2.5/1M output tokens, low-cost sufficient that engineering groups can route easy, high-volume subagent requests to a minimal pondering degree and reserve increased pondering ranges for multi-step work.

On Google’s GDM-MRCR v2 long-context take a look at, Gemini 3.5 Flash-Lite recorded a 72.2 % success price towards 60.1 % for its predecessor, and its GDPval-AA v2 rating almost doubled, from 642 to 1140. The mannequin carries the similar native computer-use instrument as 3.6 Flash.

Individually, Google says Gemini 3.5 Professional stays in companion testing forward of a full launch, and pre-training for the subsequent Gemini 4 structure is already underway.

Gemini 3.5 Flash Cyber: A restricted mannequin for patching code

Automated vulnerability scanners now floor flaws sooner than most safety groups can patch them, and that hole is the place Google positions Gemini 3.5 Flash Cyber.

The mannequin is constructed to validate and remediate code vulnerabilities, and Google studies efficiency on the CyberGym benchmark aggressive with frontier fashions (although it hasn’t made these figures public in the similar element as its consumer-facing releases.)

Distribution stays restricted to governments and vetted companions by way of a pilot programme, a limitation Google frames as a safeguard towards the mannequin producing exploit code for offensive use.

Inside Google’s CodeMender safety agent, a number of cases of three.5 Flash Cyber run in parallel, cross-checking each other’s findings before producing a single remediation report a human reviewer indicators off on.

Engineering groups looking for to combine these new fashions can entry them by way of the Gemini API by way of Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Customers also can entry the new fashions in the Gemini app and three.5 Flash-Lite is additionally rolling out in Google Search.

See additionally: Bristol Myers Squibb buys Nvidia AI system for drug discovery

Banner for the AI & Big Data Expo event series.

Need to be taught extra about AI and massive information from trade leaders? Take a look at AI & Big Data Expo going down in Amsterdam, California, and London. The great occasion is a part of TechEx and is co-located with different main know-how occasions together with the Cyber Security & Cloud Expo. Click on here for extra information.

AI Information is powered by TechForge Media. Discover different upcoming enterprise know-how occasions and webinars here.




Disclaimer: This article is sourced from external platforms. OverBeta has not independently verified the information. Readers are advised to verify details before relying on them.

0
Show Comments (0) Hide Comments (0)
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments

Stay Updated!

Subscribe to get the latest blog posts, news, and updates delivered straight to your inbox.