AI AI Toolkit
AI Newsai-models

Pro 正式版上线,Agent 能力大幅增强

DeepSeek:API 更新日志2026-08-13T11:16:42.199Z

Key Highlights

The official version of DeepSeek-V4-Pro is now live simultaneously on the app, the web client, and the API, and can be invoked by setting the model name to deepseek-v4-pro. Simply put, this is a "graduation" release: where earlier access may have been in grayscale or preview, now everyone can use it stably. The company emphasizes that its agent capability is significantly enhanced and backs the claim with two sets of hard benchmark numbers. The synchronized launch across three surfaces also means there is no feature gap between what a casual user sees and what a developer can build on. By moving from preview to general availability, DeepSeek signals confidence in the model's stability under production load. The release also tightens the linkage between the company's consumer products and its developer platform, a unification that simplifies adoption for teams that want one model across every touchpoint they operate.

What Happened and How It Worked

On agent-related benchmarks, V4-Pro's HLE (no tools / with tools) reaches 42.7 and 60.0 respectively, and Terminal Bench 2.1 stands at 87.9. HLE is a difficult set of problems that tests a model's deep, open-ended research ability, and the large jump when tools are allowed shows that the model's ceiling rises noticeably once it "knows how to use tools." Terminal Bench measures the ability to complete multi-step tasks inside a real command-line environment, and 87.9 is a rather high level. Together the two scores describe a model that is better at doing than at merely discussing. The with-tools improvement is especially relevant, because real agent work almost always involves external tools rather than pure reasoning. The numbers give buyers a concrete basis to compare V4-Pro against rivals without relying on marketing language alone, and they anchor the launch in measurable behavior rather than vibes.

Technical Details

From the numbers, V4-Pro's improvement lands mainly in "execution" rather than "chitchat." The no-tools HLE of 42.7 is already decent, but jumping to 60.0 with tools indicates it runs the retrieve-call-verify chain more smoothly. The Terminal Bench 87.9 shows it can stably run commands in a shell environment, read outputs, and then decide, possessing genuine "hands-on" ability rather than only giving advice. The pattern suggests the training and post-training work focused on tool use and environment interaction, the exact skills that separate a chatbot from an agent. The high terminal score also implies reliable long-horizon behavior, since multi-step shell tasks punish models that lose track of state. Practically, this means V4-Pro can be trusted with workflows that span many tool calls without constant human rescuing, which is the real test of an agent in production rather than in a demo.

Comparison with Competitors

On the Pro-tier agent track, Claude, GPT, and others each have their strengths. DeepSeek-V4-Pro builds trust through an open-source ecosystem, three-end synchronization, and explicit benchmarks. Simply put, it wants to prove that a domestic model can now compete on the same stage as the first tier when it comes to "completing complex tasks," and by launching on the API at the same time, developers can connect immediately. The transparency of publishing concrete scores also invites direct comparison, a confident posture for a model entering general availability. Where some competitors keep benchmark results vague, DeepSeek's explicit numbers lower the evaluation cost for skeptical buyers. The open ecosystem additionally means third parties can verify claims independently, reinforcing credibility in a market where vendor-reported scores are often met with justified suspicion.

Industry Impact and Use Cases

For teams building autonomous research assistants, automated operations, and code agents, V4-Pro is a new high-cost-performance option. Three-end synchronization means the product, the web client, and the backend can share the same capability. As agent ability becomes the decisive factor in large-model competition, the official release of V4-Pro will further heat up the main thread of "whose execution is steadier." The practical takeaway is that buyers now have another credible candidate when selecting a model to power long-running, tool-using workflows. The combination of strong tool use and broad availability also lowers the barrier for smaller teams to deploy serious agents. As the market matures, execution reliability rather than raw chat quality will likely decide enterprise adoption, and V4-Pro is positioned squarely on that axis, offering a domestically sourced alternative to the usual closed incumbents. The explicit publication of both no-tools and with-tools scores also helps set realistic expectations, showing where the model stands with and without external help. That honesty is useful in a market where agents are increasingly judged by what they can actually accomplish rather than by chat benchmarks alone. For operations teams, the Terminal Bench result is the more telling number, because it measures behavior inside a real environment where a wrong command can break state, and a score near 88 suggests the model recovers and proceeds rather than stalling. Combined with three-end availability, V4-Pro is positioned less as a research curiosity and more as dependable infrastructure for teams that run agents around the clock.