Anthropic announced that Computer Use, the Skills API and the Files API are now generally available on the Claude Platform, and introduced a new browser automation tool that lets agents operate software, call team skills and return finished files.
AI News Daily
Daily tracking of global AI industry updates, curated important news and in-depth analysis.
How Anthropic teaches people to use AI
Anthropic launched Claude Academy, an AI learning resource for millions of users worldwide that helps them use AI safely and effectively. The curriculum draws on its internal employee training methods, including the 4D AI Fluency Framework and an ever-boarding continuous learning program, and stresses problem-centered learning that builds durable mindsets rather than specific button clicks.
Hugging Face released DSpark draft-model checkpoints for three LFM2.5 models. Using speculative decoding, they raise high-end GPU throughput by up to 3.18x and on-device throughput by up to 2.87x without changing output quality. The draft models have about 300M parameters, cut average latency by 57% in LFM2.5-2.6B agent (tool-calling) scenarios, are open-source, and support llama.cpp and SGLang.
Mistral Launches Agentic Search: Multi-Step Retrieval Boosts Complex Document Query Accuracy
Mistral released Agentic Search, which uses a multi-step retrieval loop with five tools—search, open, navigate, read, and grep—to let models find, locate, and verify information across long documents and multiple sources.
AlloyDB's ScaNN index now supports over ten billion vectors through a new four-level tree architecture in preview, reducing query complexity from O(N^1/2) to O(N^1/4), with p95 latency under 51 milliseconds and 95 percent recall at that scale.
Anthropic released a Claude Code usage guide for startups, based on research with over a dozen high-growth companies, summarizing five rules: everyone can ship, automate the tedious, trust but verify, build to refactor, and prototype, use yourself, productize.
Alibaba Launches Qwen-UI Agent, Focused on Making the Model Truly "Know How to Use" Every Screen
Alibaba has officially launched Qwen-UI Agent, a real-world-centric GUI agent foundation model covering mobile, desktop, web, and DeepSearch environments.
OpenAI CFO Sarah Friar told employees at an all-hands that the company will go public no later than 2027, possibly earlier if business keeps improving. OpenAI secretly filed its IPO prospectus in June; overall annualized revenue grew 35% this quarter and enterprise revenue 50%, with weekly active users of AI coding and office products surpassing 20 million.
A 5-second 480p video generated entirely on a Mac in 30 seconds, with no CUDA, no cloud, and only 3.9 GiB of memory used. FastMetal brings the FastWan-QAD series to Apple Silicon, running DiT, the DMD sampler, and the decoder on Metal via MLX with default INT8. Three models: 1.3B for 480p, 5B for 720p, 14B for quality.
Liquid AI released four Q4_0 GGUF checkpoints - LFM2.5-230M, 350M, 1.2B-Instruct, and 2.6B - trained with QAD (quantization-aware distillation, a "slimming" technique where a smaller model is taught by a larger one). They recover 97% of the average BF16 accuracy loss while keeping native Q4_0 memory and speed.
GLM-5.3 API is live today, excelling at complex coding, defensive cybersecurity, and long-horizon tasks. It scores 60 on the Artificial Analysis Intelligence Index, on par with undisclosed internal flagship models from vendors like Anthropic and OpenAI, and ties with Kimi K3 as the top open-weight model. Smaller and cheaper, API pricing matches GLM-5.2; weights open-source next Friday.
Grok Build is now available to all plan tiers across web, iOS, and Android. Users describe an app, game, or site, and Grok generates a runnable version in chat. Since its July Early Beta, it added publishing and sharing, X integration, and Grok model calls; published apps get grok.me links with custom domains and GitHub export.
OpenRouter announced it is merging with Stripe to accelerate global economic growth. OpenRouter processes over 10 trillion tokens daily from 400+ models, serving 10M+ developers and companies, with inference volume growing at least 10x/year. Post-merger it keeps its name, mission, products, and user-first routing; deal closes in coming weeks.
Mojo is now officially open-sourced under the Apache 2.0 license (with an LLVM exception). The compiler, toolchain, and all source code have been published to Modular's GitHub repository. Mojo reached its 1.0 milestone (source-stable) just last week, and this open-sourcing covers the entire compiler and toolchain. Compiler-related contributions are not yet accepted but are planned to open by the end of the year; the standard library has accepted community contributions since 2024.
Claude can now send emails in Gmail and manage files in Google Drive. Ask Claude to reply to an email thread and it will draft and send the response. You stay in control of when your approval is required. Connect by selecting Gmail or Google Drive from the connector menu to try it. Available on all paid plans.
Citing the OpenAI–Hugging Face incident and the forthcoming Astra model potentially crossing the 'critical cyber security capability' threshold under its Preparedness Framework, OpenAI has temporarily slowed model scaling. This includes pausing two weeks of trial-and-error training on its latest deployed model and shelving its largest frontier RL run. The company has tightened research-environment security, requiring the strictest protections for Astra and cyber-related workloads, expanded process-based monitoring of step-by-step reasoning, and adopted multi-stage activation classifiers.
OpenAI launched ChatGPT for Teens, automatically enabled for users aged 13-17, with stronger safety protections and parental controls. It adds Study Mode, responsible homework reminders, quizzes and learning visualizations, plus configurable default study-time windows that guide teens to solve problems step by step rather than receive direct answers. OpenAI also announced a partnership with Common Sense Media to help teens understand, question, and creatively use AI.
This article demonstrates how to evaluate agent skills with the open-source framework Inspect AI and Harbor, and how to visualize and analyze results with Google Sheets and Data Studio.
Anthropic has temporarily slowed its model expansion, including pausing two weeks of reinforcement-learning training for its newest model, citing the OpenAI-Hugging Face incident and the upcoming Astra model potentially crossing a "critical cybersecurity capability threshold." The company tightened research-environment security with workload and network isolation and continuous safety testing.
Sentence Transformers v6.0 adds MultiVectorEncoder, supporting ColBERT-style multi-vector models
Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, which can directly load PyLate, Stanford-NLP ColBERT, and colpali-engine checkpoints for ColBERT-style late-interaction retrieval.
Inspired by a Reddit post about making Claude truly start thinking, the author introduces the 'steelman' concept from logic and devises a 'bidirectional steelman prompt for AI.' Through four steps — restating the real problem, strengthening both sides' arguments, finding the key variables, and forcing a clear judgment — it aims to avoid model sycophancy and help users find the most essential answer. The author demonstrates it with a real case of choosing a company anniversary date.
Google open-sourced a zero-trust customer-service and returns agent example built on ADK and Gemini, demonstrating how to defend against prompt-injection attacks. The architecture enforces three hard security layers outside the model context: hardware-backed cryptographic signing makes database writes non-repudiable, a gVisor sandbox isolates dynamic code execution, and a deterministic semantic gateway validates business logic. The system prompt is only a soft constraint and can never serve as a security boundary.
Cursor has opened an early-access version of Origin code hosting to all paid-plan users, offering repositories, pull requests, code browsing, and GitHub sync. Users can create repos under the cursor.com/codebase/ prefix or sync GitHub repos to Origin with two-way comment and review sync. Integrations with Vercel, Depot, and Buildkite are available, and agent features are coming soon.
Through planting a tracking device inside a rare book, 404 Media revealed for the first time Amazon's undisclosed book-buying operation: it bulk-purchases large numbers of books, scans them as material for AI learning, then destroys them. Tracking showed the books ended up at one of Amazon's AI training facilities.
A how-to guide for users who want to reduce intrusive AI in their tech environment, covering Adobe Acrobat, Android/Gemini, Apple Intelligence, Chrome, Edge, Firefox, DuckDuckGo, Google Workspace, Slack, WhatsApp, and Windows 11/Copilot.
Anthropic's CI engineers built an on-call agent with Claude to act as first responder for CI/CD failures. After an incident, Claude posts a first evidence-based analysis in a median of 14 minutes, and in the fastest case verified the fix and confirmed error rates returned to baseline within 3 minutes. The solution works through Slack channels, Datadog or Grafana tool access, and GitHub skill files. Anthropic has published a general setup kit for other teams to deploy.
Jensen Huang announced that NVIDIA and SB Energy are partnering to lock in LPS capacity at the PORTS-Pike technology park in Ohio, reserved exclusively for NVIDIA AI factories, with OpenAI as the tenant. The initial deployment is expected to provide about 4.25 gigawatts of AI-factory capacity, with roughly 1.5 million NVIDIA high-performance computing chips per generation, corresponding to $150 billion to $200 billion in revenue. OpenAI has committed to roughly 12 gigawatts of NVIDIA compute by 2030, expandable to 16 gigawatts, for a total opportunity of about $600 billion.
NVIDIA announced a partnership with SB Energy to lock in the power capacity (LPS) at the PORTS-Pike technology park in Ohio for the exclusive deployment of NVIDIA compute capability, with OpenAI becoming the tenant.
Unitree announced its shares will list on the STAR Market on August 19, 2026, at an offering price of 150.80 yuan per share, implying a market cap of about 60.993 billion yuan and expected proceeds of roughly 6.099 billion yuan. The company's revenue was 159 million, 393 million, and 1.699 billion yuan in 2023 to 2025, with net profit of -11.145 million, 95.4747 million, and 278 million yuan, making it one of the few profitable high-performance general-purpose robot companies globally.
Ant Group's inclusionAI open-sourced ConceptEdit, a pipeline for generating image-editing data based on concept scaling and dense supervision. It builds large-scale, taxonomy-based image-editing datasets through a three-stage flow, and offers single-concept and multi-concept variants. The project uses the MIT license and supports resume-from-checkpoint, requiring an OpenAI-compatible VLM endpoint and a local FLUX checkpoint.
After the OpenAI-Hugging Face incident, OpenAI reflected that it had underestimated the real cyber-attack capability of its models and is hardening security through four pillars: using Codex to verify code vulnerabilities, using agents to triage security alerts first, continuously enumerating attack paths, and opening network capabilities only to trusted defenders. The post demonstrates that ChatGPT Work (based on GPT-5.6 Sol) found 13 issues in a personal website in 15 minutes and fixed them within an hour.
OpenRouter launched an Activity dashboard and a beta Analytics API that let you view spend, token volume, and cache-hit rate by agent, model, and request, with drill-down to individual request logs.
Alibaba's Qwen lab has released Qwen 3.8 27B, a 27-billion-parameter vision-language model under the Apache 2 license. Official benchmarks show it surpasses the previous Qwen 3.6 27B and the internally developed, undisclosed Qwen 3.7-Plus.
Cursor Officially Acquired by SpaceX
Cursor has been formally acquired by SpaceX, completing a process that began in April. By tapping SpaceX’s massive high-performance compute cluster, Cursor aims to build stronger, lower-cost models and pass that power to customers at lower prices. Grok 4.6, released this Wednesday, is an early sign of the collaboration.
Gemini 3.7 Flash is now available to Pro and Ultra users inside the Gemini chat. The update improves reasoning and accuracy on multi-step tasks, such as intelligently consolidating dozens of files and emails into a single master document. Gemini Spark also now runs on 3.7 Flash, improving tool calls against Google Workspace apps to make the personal AI assistant more precise.
Qwen has open-sourced the Qwen3.8 model series. Qwen3.8-27B is a native multimodal (image/voice/text) dense model whose 27B parameters already reach the level of Qwen3.7-Plus, with native 262K context extendable to 1M tokens via YaRN, under the Apache 2.0 license. The Max-class Qwen3.8-2.4T-A95B open weights were released at the same time.
OpenAI and Anthropic are locked in an API price war as Chinese AI vendors such as DeepSeek, Qwen, and Zhipu erode their cost advantage. Falling token prices and the rise of cheap open-source Chinese models are squeezing Western labs' margins and forcing aggressive repricing of their model tiers.
Cursor Officially Acquired by SpaceX
Cursor has been formally acquired by SpaceX, completing a process that began in April. After the merger, Cursor gains access to the world's largest GPU cluster to build stronger models that run at lower cost, letting it offer more powerful models to customers at lower prices. Grok 4.6, released this Wednesday, is an early result of the collaboration.
A 280B-Parameter Lightweight Model Focused on Long-Horizon Agents and Multimodal Reasoning
Xiaohongshu Technology has open-sourced dots3-note Preview, the lightest model in the dots3 family. With 280B total parameters and 16B activated, it supports 512K context and understands text, vision, and voice, optimized for complex reasoning and long-horizon agent tasks.
Zhipu has released GLM-5.3, built on the same base as GLM-5.2 but raising the intelligence ceiling through extreme post-training scaling. Coding improved 50% over the previous generation, ranking first among open-source models on public benchmarks like Terminal Bench 3.0 and approaching Claude Code. On security tasks it matches Mythos 5 in white-box code review and scores 84.5% on CyberGym. Weights open in two weeks; available now on ZCode and AutoClaw.
DeepSeek-V4-Pro-0813 is now live on SiliconFlow with Day-0 support, offering a 1M context window and low/high/max inference-intensity tiers, with more focus on coding, tool calling, and agent workflows, still under the MIT license. Pricing is input $1.32/M, output $3.96/M, cache-hit $0.44/M. The sibling DeepSeek-V4-Flash-0731 targets everyday production scenarios that value speed and cost.
Ant Lingma and the ASystem team collaborated to run a complete single-machine agentic RL post-training loop on a DGX Spark, using Ling-3.0-tiny and AReno. With tic-tac-toe as the minimal validation task and the GSPO algorithm trained for 400 steps, rollout/rewards_mean rose from about -0.5 to 0.4, response_len dropped to about 850 tokens, and tool calls and action choices stabilized.
From January to August 2026, public model repositories on Hugging Face grew from 2.43 million to 2.96 million, yet 85.6% of models were downloaded fewer than 200 times, and 1.5% of repos captured 99.2% of downloads. Chinese labs' largest monthly open-source model ranged from 754B to 2.78 trillion parameters, while US labs stayed below 130B in five of seven months. AMD and NVIDIA each released over 200 new model repos, the most prolific open-source publishers.
A beginner-friendly walkthrough of the DeepSeek Harness, an agent framework that orchestrates models, tools, skills, and loops into reusable workflows. The piece covers what the Harness is, its core concepts, how to get started, and who it is for, using official examples and framework-level descriptions without fabricating specific APIs or version numbers.
Boris Cherny tried having Claude take over the routine maintenance of his app, running crash fuzzing, duplicate code unification, dead code removal and other daily tasks through a Slack channel. Over several weeks it automatically opened 388 PRs, of which 180 were merged after Claude Code Review and human review. Claude usually gets changes right on the first try, and when it errs, routine adjustments the next day improve it.
Google DeepMind released Gemini 3.7 Flash, only three weeks after 3.6 Flash, focused on coding and agentic tasks, with input and output prices of $0.75 and $3.75 per million tokens respectively, half of the original 3.6 Flash.
Google Sheets has launched Sheets canvas, built on Gemini. With a natural-language prompt, users can turn tabular data into interactive dashboards, study trackers, seating charts, and other "mini-apps" without writing code.
Anthropic will embed watermarks into text generated by future Claude models to estimate the likelihood that a passage was written by Claude, a change made to comply with the EU AI Act. The approach is based on Google DeepMind's SynthID-Text and has no practical impact on output quality, creativity, or readability; readers cannot distinguish watermarked text, and no extra tokens or cost are added.
Alibaba open-sourced Qwen3.8-2.4T-A95B, and SiliconFlow has provided Day-0 support. The model has 2.4T parameters with 95B active parameters, focused on autonomous coding, deep research and end-to-end agent execution. API pricing is $2.00 per million input tokens, $6.00 per million output tokens, and $0.25 per million cached input tokens.
DeepSeek Harness v0.1 is now available as a developer preview and open-sourced under the MIT license. The agent framework is built on the Cordis meta-framework, with a core design of "everything is a plugin," where models, tools, skills, sessions, sandboxes, file systems, loops, orchestration and UI can all be freely combined, replaced and extended.
Cloud Agent Startup Speed Boosted to 3x
Cursor introduced the builds feature, which continuously prepares ready-to-go copies of the development environment in the background, so cloud agents no longer need to build from scratch at startup, with response speed up to 3x faster. Internal environment startup is 10x faster and first token generation is 3x faster; agents always start from the most recent successful build, and dependency updates or install-script errors will not affect running. From August 17, builds are enabled by default for all environments at no extra cost.
DeepSeek-V4-Pro official version is now live simultaneously on the app, web, and API, and can be used by setting the model name to deepseek-v4-pro. Its agent capability is significantly improved, with HLE (no tools / with tools) reaching 42.7 / 60.0, and Terminal Bench 2.1 at 87.9.
The GPT-5.6 model family delivers frontier-level agent performance at lower cost, and adds API capabilities such as reasoning persistence, native multi-agent orchestration, and programmatic tool calling. On ARC-AGI-3, with retained reasoning and compression enabled, Sol score jumped from 13.3% to 38.3% while output tokens dropped by about 6x. Luna matched GPT-5.5 on BrowseComp at 84.04% versus 84.36%, with cost falling from $33.27 to $1.33.
Xiaohongshu (RED) dots team open-sourced dots.tts, a 2-billion-parameter fully continuous end-to-end autoregressive speech synthesis model, achieving the best average content accuracy and average speaker similarity on the three subsets of Seed-TTS-Eval.
WorkBuddy updated with a remote control feature that connects PC, app, and mini-program, letting the phone sync in real time the tasks, conversations, workspace, and artifacts from the computer, supporting one phone connecting to multiple computers and switching at any time. The app needs to be upgraded to 1.2.0 or above and the computer client to 5.3.8 or above, with no QR code needed to connect. The update also adds a library (My Documents and Team Space), Markdown multi-user collaborative editing, an AI-native review mode, and generating the library content into publishable HTML websites.
DeepSeek V4 Pro and xAI's Grok 4.6 were released within a two-hour window, with roughly 1.6T and 1.5T parameters respectively, both nearing the coding experience of Claude Code.
AutoGPT maintainers stopped trusting agents to read docs and instead baked rules into AGENTS.md and skill files, gating agent pull requests with mandatory templates, CI coverage, and a CLA signature that acts as a human detector.
Microsoft has released MAI-Thinking-1, its first self-built reasoning model, now available on Microsoft Foundry and announced by AI CEO Mustafa Suleyman.
MiniMax launched Music 3.0, a new-generation music generation model that can complete the composition, arrangement, performance, and production of an entire song at once from a creative concept and optional lyrics, supporting up to five minutes in length.
After writing an RLHF textbook, Nathan Lambert argues that models have stalled at long-form nonfiction writing while soaring at coding and math, making them useful for editing but not for authoring whole books.
Anthropic has upgraded the Claude Chrome extension sidebar into a full Claude Cowork session that is saved to history, syncs across desktop, web, and mobile, and runs skills and connectors directly in the browser.
Alibaba Qwen Team Open-Sources Qwen3.8-2.4T-A95B Weights, First Qwen-Max-Level Model Released
Alibaba's Qwen team open-sourced the Qwen3.8-2.4T-A95B model weights, its first Qwen-Max-level model released for free use, with 2.4T parameters, 95B active per token, and a context expandable beyond one million tokens.
Meta AI Superintelligence Labs has released Muse Glimmer, a 30B dense text-and-image model under Apache 2.0, positioned as a reliable local agent and scoring 75.5 on MCP Atlas and 51.2 on SWE-Bench Pro.
LTX-2.5 produces a 10-second 720p video in 6.8 seconds on two GB200 GPUs with native ComfyUI support, and is free for organizations under 10 million USD ARR.
Investigators found that Research Gold, which advertises 100% human-written medical content, uses AI-generated fake PhD reviewers, steals real experts' identities, and answers calls and chats with an AI that insists it is human.
The article lays out a 12-step practical workflow for absolute beginners to get started with AI in half a day: prepare a computer with at least 16GB of RAM, subscribe to ChatGPT and install Codex or use WorkBuddy, brief the AI on tasks via voice input using a [background, pain point, need] framework, let it clarify requirements through Socratic questioning, then feed files so the AI completes the work directly, and finally distill the experience into a reusable Skill. It suggests choosing GPT-5.6 Sol as the top model for Codex and Kimi K3 for WorkBuddy.
Cursor and SpaceXAI launched Grok 4.6, boosting long-running agent and interactive vision skills, matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index.
xAI released Grok 4.6, upgrading from 4.5 with stronger long-running agent and visual abilities, matching GPT-5.6 Sol on the nine-benchmark Intelligence Index.
OpenRouter's live leaderboard shows that raising search budget from 1 to 25 rounds nearly doubles BrowseComp scores at 2.5-7x cost, that model choice beats search engine choice, and that high-failure tasks should use lower search depth.
Modular's Mojo language has hit 1.0, offering a stable production-ready base; since 2023 it grew into a general-purpose language with nearly 200 contributors, over 1,100 PRs, and more than 200,000 changed lines after open-sourcing its standard library.
Both OpenAI's and Google's chatbots have crossed the 1-billion-user threshold. In an August 6 blog post, OpenAI disclosed that ChatGPT has surpassed 1 billion monthly active users, while Google CEO Sundar Pichai announced that Gemini has reached 1 billion monthly actives, making it the fastest-growing product in the company's history. OpenAI said ChatGPT hit 1 billion weekly active users back in July, and Gemini had 750 million monthly actives in February, showing an even stronger growth trajectory.
Full cast, full tracks, generated one at a time. Seedance 2.5 is now live on Runway, supporting 50 unique character references and clips of up to 30 seconds that are synchronized with music.
The author used mitmproxy to intercept GitHub Copilot inside VS Code via a man-in-the-middle proxy, reverse-engineering its network traffic and internal architecture. The article notes that these AI apps are mostly built on Electron and share similar network stacks, so the findings transfer to other similar applications. Through this, the author reveals Copilot's runtime behavior and shares concrete steps for configuring the proxy.
You can now keep the work of other AI agents in sync with ChatGPT Work and Codex. You can import projects, chat histories, skills, and plugins, view import history, and choose to enable automatic updates in settings. It is now available in the ChatGPT desktop app.
Researchers Discover API Vulnerability That Can Read Encrypted Reasoning of Models Like ChatGPT
Alexander Panfilov's team found a vulnerability in the APIs of major AI providers including OpenAI, Anthropic, and Google that can expose the encrypted thinking process of reasoning models. Scanning about 7,000 public sessions revealed 62 API keys, 33 emails, and 33 passwords. Through jailbreaking, Anthropic's Haiku 4.5 could transcribe Opus 4.8's raw reasoning verbatim; decoding 10,000 reasoning traces cost about $720 in API fees.
More than 1 billion people now use @Geminiapp every month to spark new ideas and get work done. This is the fastest-growing product we've ever had, and our 14th product to reach the 1-billion-user milestone. Thanks to @JoshWoodward and the entire Gemini team, and to everyone who has come along with us—more exciting things ahead!
Dwarkesh Patel and Ryan Greenblatt, chief scientist at Redwood Research, discussed the possibility of recursive self-improvement (RSI): once AI reaches top-human-expert level, it could achieve AI progress equivalent to 4–5 years within a single year, and Ryan's median expectation is automated AI R&D by 2031. They also discussed risks such as superintelligence alignment—keeping AI behavior in line with human expectations—and whether reward hacking could escalate into AI joining forces to take over the world.
Google Cloud introduced Gemini-powered AI-assisted code conversion in Database Migration Service (DMS), which can convert stored procedures, triggers, and custom functions from Oracle or SQL Server into PostgreSQL PL/pgSQL code.
Nvidia is developing a new generation of open-source AI models, Nemotron 4, with the largest model expected to have at least 1 trillion parameters, aimed at competing with the world's most advanced open-source models. Nvidia has not set a release date and final training is unfinished; employees believe the model could be ready as early as late this autumn. The move is meant to expand AI adoption through an open model ecosystem and drive demand for its GPU compute capacity.
With new cross-session messaging in Codex and Claude, a main-plus-branch conversation structure replaces handoff docs and git backups, reshaping how developers collaborate with AI on code.
SGLang announced day-zero support for NVIDIA Nemotron 3.5 Lightning, an open-source mixture-of-experts (MoE) model with 30B total parameters and 3B activated parameters, supporting context lengths up to 1M tokens. BF16 and NVFP4 weights are available on Hugging Face. The model supports three speculative decoding techniques—MTP, DFlash, and DSpark—and can plug into agent workflows through an OpenAI-compatible API.
NVIDIA released Nemotron 3.5 Lightning, a customizable open-source 30B mixture-of-experts (MoE) model designed for resident AI agents. Compared with similar open-source models, its token generation speed is up to 4x faster and task completion time is reduced by 30%. The model uses open weights, supports task-specific fine-tuning, and runs on RTX PCs, DGX Spark, and Jetson devices.
OpenAI announced that its unreleased Astra model solved 10 long-standing open math problems spanning sphere packing, error-correcting codes, and the existence of non-sofic groups, and published a 250-plus-page paper along with Lean verification results.
Ant Ling (Bailing) has open-sourced Ling-3.0-tiny, a native hybrid reasoning model with 7.9B total parameters that activates only 1.3B parameters during inference, and simultaneously releases three versions in BF16, FP8 and INT4.
ZCode, deeply optimized for GLM, launched four features today: Goal, Subagents, Remote Control and idle tasks. In the Z.ai Code Bench test, GLM-5.2 with ZCode achieved a 2.39% higher overall task pass rate than with Claude Code; ZCode's cache hit rate exceeds 98%, and with a 1.5x limited-time quota bonus, overall GLM Coding Plan usage approaches 1.8x the regular quota.
This tutorial demonstrates how to use ComfyUI as a headless inference backend to build an end-to-end MiniMax-H3 video generation workflow. By constructing the execution graph directly in Python, it supports text-to-video, first/last frame conditional generation and reference image conditional generation, and automatically selects among quality, balanced and squeeze weight configurations based on accelerator VRAM. The pipeline covers automatic model download, node schema validation, joint audio-video decoding and progress monitoring, reproducing experiments without a GUI.
Challenging the popular claim that dynamic languages save more AI model tokens than static ones, the author used GPT-5.6 Sol to have an AI agent implement a zstd decoder in a real test. Results showed dynamic languages performed better at medium effort, while static languages actually won at ultra effort, and earlier benchmarks had flaws such as incorrect test paths. The author argues performance on trivial tasks does not generalize to larger problems.
Anthropic is in talks with potential investors to prepare what could become the largest IPO in history, planning to list officially in September or early October. The company is valued at up to $965 billion with annualized revenue already exceeding $47 billion, and is downplaying the competitive impact from Chinese AI firms. Anthropic also plans to expand AI applications in healthcare and biology, but has not yet announced specific IPO pricing.
Based on Xiaowei, WeChat launched internal tests of Moments AI writing-assist and AI commenting: the former generates three Moments captions from an image and written text, while the latter lets you long-press text to generate a comment or send a quick reply. The author argues these two features place AI at the core of social interaction and may encourage AI-generated content, undermining the 'record beautiful life' tone Moments has held since 2012. On the public account side, Xiaowei also stays pinned at the top and auto-summarizes frequently read articles, which the author worries will in turn reshape creators' content production.
While online-services valuations are under broad pressure overall, each niche segment has produced an AI leader: CrowdStrike leads security at 34.4x forward revenue (versus a 3.9x median), Cloudflare 32.6x vs 17.5x, and Shopify 11.3x vs 1.4x.
Databricks explains how to let Genie agents run on both structured data and documents at once without sacrificing governance. The article explores how building agents that automate simple business tasks is easy, but fusing two data sources under a unified governance framework and executing queries securely and compliantly still requires solving key problems such as data permissions, lineage tracking, and policy control.
Nvidia announced a partnership with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish an independent financing platform mobilizing over $500 billion in third-party capital to support AI infrastructure buildout.
We recently made auto mode the default option in Claude Code, meaning you no longer need to approve every action. But what determines whether an action can run safely? Here's how it works:
The Firestore database of AI meeting-recording platform tl;dv lacked tenant isolation, letting any authenticated user query all 181,000 meeting records—covering 84,312 users and 35,003 domains, including meetings of 23 governments and multiple universities. About 1,000 meetings in recording state exposed joinable meeting IDs, which researchers used to break into live calls of Malaysia's education ministry and a US university startup team. The flaw remained unpatched six months after being reported in January 2026, with over 1,000 recordings also left public.
a16z data shows the best score of computer-use agents on the OSWorld-Verified benchmark has risen from 42% a year ago to 85%, surpassing the roughly 72% level of human testers, with Claude Fable 5 leading at 85%.
SGLang, in collaboration with Meta Superintelligence Labs, provides day-zero support for Muse Glimmer, a 30B-parameter multimodal model with a 128k+ token context window.
Meta introduced Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. It outperforms leading same-size models on key agent benchmarks and standardized evaluations, and is designed to run entirely on consumer hardware such as Macs or PCs with high-performance GPUs, released under the permissive Apache 2.0 license.
Scale AI announced it will soon release open-weight versions of Muse Spark 1.2, and is releasing Muse Glimmer—a 30B-parameter agent model under the Apache 2.0 license. Muse Glimmer runs on 24GB of VRAM without losing agent reliability.
I believe everyone should be able to use superintelligence. I've written a long essay laying out Meta's philosophy and values for building a positive future for all. http://meta.com/thefutureisforeveryone
OpenAI released GPT-5.6-Cyber, a cybersecurity-specific model available via Daybreak Red, for authorized vulnerability research, exploit validation, and security testing. The model addresses the shrinking cyber defense window and gives security researchers a dedicated toolset.