Weekly AI Review β August 16, 2026
π Model & Product Releases
- Google prices Gemini 3.7 Flash at half its predecessor β Google released the coding and agent model on August 13 with a one-million-token input window, a 64,000-token output limit and multimodal input. Introductory API pricing is $0.75 per million input tokens and $3.75 per million output through the end of 2026, with standard rates of $1.50 and $7.50 from January 1, 2027. Google reports DeepSWE v1.1 rising from 49.0 to 65.3 percent against Gemini 3.6 Flash.
- Meta returns to open weights with 30-billion-parameter Muse Glimmer β The model shipped on August 10 under an Apache 2.0 licence, distilled from Meta’s larger Muse Spark and compressed to roughly 4-bit precision so it fits under 20 GB on a single consumer GPU. Weights are on Hugging Face, with llama.cpp, MLX, vLLM and SGLang support. Meta describes it as its first open-weights release since Llama 4.
- Z.ai ships GLM-5.3 but staggers the open weights β The model, released on August 14, keeps the 744-billion-parameter base of GLM-5.2 and takes its reported gains from extended post-training alone. Z.ai reports 84.5 percent on CyberGym against 83.8 for Claude Mythos 5 and 83.6 for GPT-5.6 Sol; none of those figures has been independently reproduced. Weights and wider API access follow in stages after roughly two weeks of safety evaluation.
- OpenAI previews an Ultrafast tier running on Cerebras hardware β The preview announced on August 13 runs GPT-5.6 Sol on Cerebras wafer-scale engines, which hold model weights in 44 GB of on-chip SRAM. OpenAI and Cerebras report up to 14 times standard speed and up to 750 output tokens per second. Access is a limited preview for selected customers, with no pricing and no general availability date.
π¬ Research Highlights
- Small recurrent model claims an ARC-AGI cost-efficiency record β A preprint posted on August 10 (arXiv 2608.09888) describes a 150-million-parameter model that updates a recurrent memory from context, then reasons iteratively in latent space without emitting intermediate text. The authors report 29.5 percent pass@2 on ARC-AGI-1 at a computed inference cost of $0.0007 per task, which they call a new state of the art in benchmark cost efficiency.
- Agent pipeline generates research papers with integrity checks β A paper posted on August 12 (arXiv 2608.11924) splits research-paper generation into thirteen composable skills inside a coding assistant. The authors report 99.5 percent citation validity, 96.4 percent figure editability, and fabrication detection rising from 14 percent for a single-pass draft to 92 percent with the full integrity and review stack. They name a self-refutation loop as a failure mode.
- Mixture-of-LoRA design targets continual learning on a frozen base β A Mind Lab preprint posted on August 10 (arXiv 2608.09819) freezes a 744-billion-parameter GLM-5.2 base, attaches four specialist LoRA adapters and selects one per user turn. Specialists can be added without retraining the base. The authors state that compounding gains from continual learning and collective intelligence remain open questions.
ποΈ Infrastructure & Compute
- Nvidia enlists six firms for $500 billion in chip financing β Announced around Nvidia’s GTC conference on August 10 and 11, the arrangement pairs the company with Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs and KKR to source up to $500 billion of third-party financing for AI infrastructure purchases. Jensen Huang told CNBC he approached only those six firms and none declined. Bank of America and Morgan Stanley backed the structure; Wells Fargo and Mizuho questioned its sustainability.
- IBM and Together AI commit $240 million to a B300 cluster β The multi-year agreement announced on August 11 puts Nvidia HGX B300 systems and Spectrum-X Ethernet networking on IBM Cloud, in what the companies call the first dedicated large-scale B300 inference cluster there. Availability is targeted for the first quarter of 2027. Together AI, valued at $8.3 billion after its Series C, says it serves 400 trillion tokens a month.
- Reuters reports Nvidia training a trillion-parameter open model β Reuters said on August 11, citing unnamed employees, that Nvidia is training Nemotron 4, an open-model family whose flagship would exceed one trillion parameters, roughly twice the size of Nemotron 3 Ultra, and could be ready as early as late autumn. The report carries no benchmarks, no licence terms and no training-data disclosure, and Nvidia has not confirmed it. Nvidia did release Nemotron 3.5 Lightning and the NeMo Switchyard routing library the same day.
πΌ Industry & Funding
- Anthropic in talks to buy Decart for $6 billion β Bloomberg reported on August 13 that Anthropic is negotiating what would be its largest known acquisition, for the Israeli startup whose software raises chip efficiency in training and inference. Decart raised $300 million in May at a valuation near $4 billion, led by Radical Ventures with Nvidia and Adobe Ventures taking part. Neither company has confirmed a deal.
- Lovable raises $400 million at a $13.3 billion valuation β The Swedish app-building company announced its Series C on August 12, roughly doubling its previous valuation. It put annualised revenue near $600 million.
- Google says the Gemini app passed one billion monthly users β Sundar Pichai announced the milestone on August 11, calling Gemini the fastest-growing product in the company’s history and the fourteenth Google product to reach a billion users. Google published no paid-subscriber figure alongside it. OpenAI reported ChatGPT crossing the same threshold weeks earlier.
- Blacksmith raises $45 million as code-validation demand grows β The San Francisco continuous-integration company announced its Series B on August 12, led by Peak XV Partners with GV and Y Combinator taking part, at a $550 million valuation against roughly $60 million less than a year earlier. It says its customer base grew from about 800 companies to more than 6,000.
βοΈ Policy & Safety
- Anthropic starts watermarking text from its models β The company said on August 11 that every model released after August 2 carries a model-level watermark across the Claude API, Claude, Claude Code and Claude Cowork, with older models to follow, and that files use the C2PA standard. Anthropic says the mark travels with copied text, may survive some editing, and evidences processing by its models rather than authorship. The move follows EU AI Act transparency rules enforceable since August 2; commentators argued the same week that paraphrasing through a second model strips such marks.
- Industry alliance drafts an AI incident reporting framework β Axios reported on August 11 that the Open Secure AI Alliance, whose members include Nvidia, Cisco and CrowdStrike, circulated a draft called the Shared AI Findings Exchange. Signatories would report cases in which an AI system reaches or exploits a third-party system without authorisation, disclose near misses, and preserve prompts, agent traces and tool calls. Government agencies would take part as non-controlling observers; the draft is voluntary and not final.
- Researchers document near-autonomous AI attack on Taiwanese agencies β The Israeli firm Dream published research on August 12 describing a four-day campaign from July 1 to July 4 against Taiwanese government systems, a nuclear safety agency, supply-chain vendors and energy companies. It counts 85 compromised accounts, more than 2,500 personnel records, seven single-sign-on client secrets and six database credentials, taken across twelve waves against 21 connected systems. Dream says the framework was built on the open-source Hermes and OpenClaw agents and that operator documentation is in simplified Chinese, but stops short of state attribution.
- California committees clear most AI bills on suspense day β The Assembly and Senate appropriations committees held their suspense-file hearings on August 13, covering roughly 29 active AI measures on chatbot safety for minors, copyright transparency and algorithmic management. Five bills were held in committee, which ends them for the session without a recorded explanation, and the remainder advanced to floor votes. Two went to the governor: AB 1651 on AI in the state bar exam and SB 928 requiring California State University instructors to be human.
π What to Watch
- GLM-5.3 weights β Z.ai says open weights and wider API access follow in stages after roughly two weeks of safety evaluation, putting them near August 28.
- Nvidia earnings β The company reports second-quarter fiscal 2027 results after the close on August 26.
- Imagen 4 shutdown β Google’s imagen-4.0 standard, fast and ultra generation endpoints stop serving on August 17, with the Gemini image models as the migration target.
Weekly AI review generated on August 16, 2026.
