Alibaba’s Qwen38-Max Claims Five Days of Autonomous AI Work — A New Front Opens in the US-China AI Race + Video

Listen to this Post

Featured ImageA New AI Rivalry Is Moving Beyond Chatbots

The artificial intelligence race is entering a more consequential phase. The competition between Chinese and American technology companies is no longer limited to who can produce the smartest chatbot, generate the best image, or write the cleanest piece of code. The next battlefield is autonomy: how long an AI system can work independently, how complex a task it can manage, and whether it can produce meaningful results without a human constantly guiding every step.

Alibaba’s Qwen team is now making a striking claim about its latest model, Qwen3.8-Max. According to the company, the model can operate independently for days, tackle complicated research and engineering problems, and continue working without continuous human supervision.

If those claims withstand independent scrutiny, the development could mark an important shift in how AI systems are evaluated.

The announcement also arrives at a politically sensitive moment. The United States and China have spent years competing over advanced semiconductors, AI infrastructure, computing power, and access to cutting-edge technology. Washington has imposed restrictions intended to limit China’s ability to acquire some of the most advanced AI chips, while Chinese companies have increasingly focused on developing domestic alternatives and improving the efficiency of their models.

Alibaba’s latest Qwen announcement therefore carries significance far beyond a single model release.

It is also a direct challenge to the idea that the most advanced autonomous AI systems will necessarily come from Silicon Valley.

Qwen3.8-Max Wants to Work Without Constant Human Supervision

Alibaba’s Qwen team describes Qwen3.8-Max as a broad improvement across coding, research, professional work, and long-horizon tasks.

The most interesting part of the announcement, however, is not a conventional benchmark score. It is the claim that the model can continue working on a difficult problem for an unusually long period of time.

According to Alibaba, one experiment involved Qwen3.8-Max working independently for approximately 125 hours, or five continuous days.

That is a very different proposition from asking an AI model a question and receiving an answer within seconds.

A long-running autonomous system has to repeatedly make decisions, inspect its own work, identify problems, change direction, test hypotheses, and continue making progress. The challenge is not simply generating text. It is maintaining a coherent strategy over an extended period.

That distinction could become one of the most important measurements of the next generation of AI.

The Five-Day Research Experiment

Alibaba says the model was given a mathematical research task involving the reconstruction of an experiment from an existing research paper.

Rather than simply summarizing the paper, Qwen3.8-Max reportedly rebuilt the experiment from scratch and then attempted to improve upon it.

This type of task is particularly interesting because it requires multiple stages of reasoning.

The model must first understand the research objective, determine what the original experiment is attempting to demonstrate, recreate the necessary methodology, work through technical obstacles, evaluate the results, and potentially identify ways to improve the experiment.

Doing that over several hours would already be impressive.

Doing it continuously over roughly five days would represent a much more ambitious test of AI autonomy.

However, the distinction between a company claiming that a model accomplished something and independent researchers verifying that accomplishment is extremely important.

At this stage,

Qwen Also Claims a Competitive Advantage Against Human Teams

Alibaba described another experiment involving a live online competition.

According to the Qwen team, the model participated in a contest alongside 526 human teams, operating under a 24-hour deadline.

Alibaba says Qwen3.8-Max finished ahead of all but 68 of those teams.

If accurate, that would place the model ahead of hundreds of participating human teams under a real-world time constraint.

The result is particularly interesting because competitive environments are different from carefully constructed benchmark tests.

A benchmark typically has a predefined structure. A competition introduces uncertainty, time pressure, strategic decisions, incomplete information, and the possibility that participants will take very different approaches to solving the same problem.

Still, the precise competition, evaluation methodology, participant capabilities, and degree of AI assistance involved are essential to understanding what the result actually means.

Without independent validation and detailed methodology, the number is impressive but should not automatically be interpreted as proof that AI has surpassed humans broadly.

Anthropic Is Making a Similar Bet on Autonomous AI

Alibaba is not alone in pushing this direction.

Anthropic has increasingly positioned its Claude ecosystem around the concept of AI systems that can take responsibility for longer and more complicated workflows.

Its coding tools are designed to allow AI to inspect software projects, plan modifications, implement changes, run tests, analyze failures, and iterate with comparatively limited human intervention.

That represents an important evolution from traditional chatbot interaction.

Instead of asking an AI to perform one isolated action, the user increasingly gives the system a goal.

The AI then decides which sequence of actions may be necessary to reach that goal.

This is the foundation of what is often called agentic AI.

From Answering Questions to Completing Jobs

The biggest change taking place in AI may not be that models are becoming better at answering questions.

It may be that they are becoming better at completing entire jobs.

A traditional chatbot waits for instructions.

An autonomous AI agent can potentially interpret an objective, create a plan, execute multiple steps, check its progress, recover from errors, and continue until it reaches a result.

That distinction has enormous economic implications.

Imagine asking an AI to investigate a technical problem rather than asking it to explain the problem.

Instead of producing a paragraph, the system could potentially search documentation, inspect code, create experiments, run tests, compare results, identify errors, revise its approach, and deliver a final report.

The longer an AI can perform those tasks reliably, the less human supervision may be required.

The Real Competition Is About Reliability

Raw intelligence is only one part of autonomous AI.

A model can be extremely capable and still be dangerous or inefficient if it makes one bad assumption after another and continues operating without anyone noticing.

The real question is therefore not simply:

How smart is the model?

It is:

How reliably can the model remain useful while operating independently?

That requires measuring error rates, recovery behavior, hallucinations, planning quality, security boundaries, resource consumption, and the ability to recognize when it does not know something.

Five days of autonomous operation sounds impressive.

Five days of autonomous operation with incorrect decisions accumulating silently would be much less impressive.

The difference between those two scenarios will define the next stage of the AI industry.

China and America Are Now Competing on the Same AI Frontier

The geopolitical dimension makes Alibaba’s announcement even more significant.

For years, the United States has maintained a substantial advantage in advanced AI infrastructure, access to cutting-edge semiconductor technology, and the ecosystem surrounding the world’s largest AI companies.

At the same time, Washington has increasingly restricted the export of advanced AI chips and related technologies to China.

The objective has been to constrain China’s ability to obtain or manufacture some of the world’s most capable AI computing systems.

But the Chinese AI industry has responded by emphasizing model efficiency, domestic semiconductor development, open models, and aggressive engineering.

Alibaba’s Qwen program is one of the most visible examples of that strategy.

The message behind the latest announcement is therefore straightforward:

China does not necessarily need to win every battle over hardware if its companies can continue improving the software and models running on the hardware they can access.

Open Weights Change the Nature of the Competition

Another important part of the Qwen strategy is openness.

Alibaba has increasingly released models whose weights can be accessed by researchers and developers, rather than keeping everything exclusively behind a proprietary API.

That creates a fundamentally different competitive dynamic.

A closed model asks users to trust the company operating it.

An open-weight model gives researchers and developers greater opportunities to inspect, test, modify, benchmark, and deploy the system themselves.

That does not automatically mean every claim becomes true.

But openness can make independent experimentation considerably easier.

If Qwen3.8-Max really demonstrates exceptional long-horizon capabilities, an open distribution strategy could allow those capabilities to spread rapidly through universities, startups, developers, enterprises, and other AI laboratories.

The Open-Model Strategy Could Become China’s Strongest Advantage

The AI race is often described as a battle between individual companies.

In reality, it may increasingly become a battle between ecosystems.

An impressive proprietary model can generate enormous revenue for its creator.

An impressive open model can become infrastructure for thousands of other organizations.

That distinction matters.

If developers build applications around a model, researchers fine-tune it, companies deploy it internally, and other laboratories improve upon it, the original model can have an impact far beyond its initial release.

China has shown a willingness to compete aggressively in this area.

Alibaba, DeepSeek, and other Chinese AI organizations have helped demonstrate that open and relatively accessible models can become powerful tools for global developers.

The United States remains extremely strong in frontier proprietary AI.

China may be attempting to build a different kind of advantage around accessibility, engineering speed, and open deployment.

Autonomous AI Could Transform Software Development

One of the clearest areas where these models could have an immediate impact is software engineering.

Modern software projects contain enormous amounts of repetitive work.

Developers spend time reading documentation, navigating unfamiliar codebases, writing tests, investigating bugs, reviewing dependencies, updating configurations, and documenting changes.

An autonomous coding agent could potentially handle large portions of that workflow.

Instead of asking an AI:

How do I fix this bug?

A developer could eventually say:

“Find the cause, reproduce it, implement the safest fix, test it, and prepare the changes for review.”

That is a fundamentally different relationship between humans and AI.

The developer becomes the supervisor rather than the person manually executing every technical step.

Research Could Become Another Major Target

Scientific and technical research may be even more important.

Researchers routinely spend significant amounts of time searching literature, reproducing experiments, processing data, writing code, checking results, and preparing drafts.

An autonomous AI system capable of handling these stages could dramatically accelerate research workflows.

But research also exposes one of the biggest weaknesses of current AI systems.

An AI can confidently produce an incorrect result.

It can misunderstand an experimental setup.

It can use inappropriate assumptions.

It can create flawed code.

And it can sometimes fail to recognize that its own output is wrong.

For that reason, autonomous research systems will require robust verification mechanisms rather than simply longer execution times.

Five Days of Work Does Not Mean Five Days of Correct Work

This distinction deserves special attention.

When a company says its model worked for 125 hours, the headline number can easily dominate the discussion.

But duration alone is not intelligence.

A human researcher who spends five days making progress is valuable.

A machine that spends five days repeatedly making the same mistake is not.

The most important measurement is therefore productive autonomy.

How much useful work was completed?

How many errors occurred?

How often did the system need to restart?

How much human intervention was quietly required?

What percentage of the final result survived independent review?

Those questions are much more important than the raw number of hours.

The Benchmark Era May Be Ending

For years, AI companies have competed heavily on standardized benchmarks.

Models were compared using mathematics, coding, reasoning, language understanding, knowledge tests, and other controlled evaluations.

Benchmarks remain useful.

But they increasingly struggle to capture what people actually want from AI systems.

A business does not necessarily care whether a model achieves a certain score on a test.

It cares whether the AI can complete a complicated workflow without wasting time or creating expensive mistakes.

A researcher cares whether the system can produce reproducible results.

A programmer cares whether it can safely modify a real codebase.

A security team cares whether an AI agent can operate without exposing credentials or creating vulnerabilities.

The next generation of benchmarks may therefore focus increasingly on long-running tasks performed inside realistic environments.

Autonomous AI Creates a New Cybersecurity Problem

Greater autonomy also creates greater risk.

An AI agent that can browse the internet, execute code, access files, interact with APIs, and make decisions for hours or days has a much larger potential attack surface than a chatbot that simply answers questions.

Prompt injection becomes more serious.

Malicious websites become more dangerous.

Compromised software dependencies become more consequential.

A single manipulated document could potentially influence an autonomous agent’s decisions.

This means AI security will increasingly have to resemble traditional cybersecurity.

Identity controls, permissions, sandboxing, monitoring, logging, network segmentation, secret management, and human approval gates will become critical components of autonomous AI systems.

The Most Powerful AI May Also Need the Strongest Restrictions

There is an uncomfortable paradox at the center of autonomous AI.

The more capable the system becomes, the more useful it is to give it access to real tools.

But the more access it receives, the more dangerous a mistake becomes.

An AI that can only generate text has limited ability to cause real-world damage.

An AI that can modify production code, access corporate databases, send emails, purchase resources, or execute commands has dramatically more power.

That means future AI development will not simply be a race toward greater autonomy.

It will also be a race toward controlled autonomy.

The winning systems may not be the ones that can operate for the longest.

They may be the ones that can operate independently while remaining observable, interruptible, auditable, and constrained.

The Human Role Is Changing Rather Than Disappearing

The rise of autonomous AI does not necessarily mean humans become irrelevant.

Instead, human work may move upward in the decision-making hierarchy.

People may increasingly define objectives, establish constraints, evaluate results, approve important decisions, and intervene when systems behave unexpectedly.

This resembles the transition from manual calculation to computers.

The computer did not eliminate the need for mathematicians.

It changed what mathematicians could accomplish.

Autonomous AI could produce a similar transformation across software engineering, research, finance, administration, design, cybersecurity, and other knowledge-intensive fields.

The Economic Impact Could Be Much Larger Than Chatbots

If AI systems become capable of working independently for many hours, their economic value could increase dramatically.

A chatbot generates an answer.

An autonomous agent potentially generates an outcome.

That difference matters to companies.

Businesses can justify paying for software that saves employees hours of repetitive work.

They may be even more willing to pay for systems that can complete entire workflows overnight.

Imagine a company assigning an AI system a technical research project at 6 p.m. and receiving a structured report, tested code, supporting evidence, and recommendations the following morning.

That is much closer to hiring digital labor than using a traditional chatbot.

The AI Labor Market Could Become More Competitive

This development could eventually change how companies think about labor.

Instead of asking whether AI can replace a particular profession, businesses may begin asking which individual tasks inside that profession can be automated.

That is a more realistic and potentially more disruptive question.

A software engineer may continue to be necessary, but the number of engineers required to maintain a particular codebase could change.

A researcher may remain essential, while AI handles literature searches and preliminary experiments.

A financial analyst may remain responsible for decisions while AI prepares models and reports.

The result could be a gradual restructuring of jobs rather than an overnight replacement of entire professions.

Why Alibaba’s Claim Matters Even Without Immediate Verification

Even if

Companies do not make these announcements in a vacuum.

Alibaba is signaling what it believes the market should measure next.

Anthropic is doing something similar.

Other AI companies are likely to follow.

The competition is moving from:

Which model answers the question better?

toward:

“Which model can take responsibility for the entire task?”

That is a much more consequential competition.

Deep Analysis: The Autonomous AI Race Has Entered a New Phase

Command: Watch the Duration

The five-day claim is important because long-horizon execution is fundamentally different from short-response reasoning.

Command: Measure the Work

The key metric should not be how long an AI remains active, but how much verified useful work it completes during that period.

Command: Verify the Experiment

Alibaba’s claims should be independently reproduced before the wider industry treats them as established performance.

Command: Separate Marketing From Evidence

Every AI company has an incentive to highlight its strongest demonstrations, making independent evaluation essential.

Command: Compare Real Tasks

The future of AI benchmarking should increasingly involve messy, realistic tasks rather than only clean laboratory tests.

Command: Track Human Intervention

An “autonomous” system is not truly autonomous if engineers repeatedly intervene behind the scenes.

Command: Examine Failure Recovery

The ability to detect and recover from mistakes may matter more than raw reasoning speed.

Command: Test Long-Horizon Memory

Agents operating for days must maintain context without gradually losing track of their objectives.

Command: Measure Goal Stability

An autonomous system must remain focused on the original objective instead of drifting toward easier but irrelevant tasks.

Command: Audit Tool Usage

Models with access to browsers, terminals, APIs, and files require much stronger oversight than ordinary chatbots.

Command: Harden the Sandbox

Autonomous AI should operate inside carefully controlled environments with restricted permissions.

Command: Protect Credentials

Long-running agents should never receive unrestricted access to sensitive credentials.

Command: Log Everything

Organizations deploying autonomous AI will need detailed records of what the system did, when it did it, and why.

Command: Build Human Checkpoints

High-impact actions should still require approval even when routine operations can be automated.

Command: Test Against Prompt Injection

The longer an agent interacts with external information, the more opportunities attackers have to manipulate it.

Command: Evaluate Reproducibility

A successful demonstration should be repeatable rather than dependent on one carefully selected example.

Command: Examine Compute Efficiency

A model that requires enormous computing resources may be less commercially useful than a slightly weaker system that is dramatically cheaper.

Command: Watch Open-Weight Adoption

Qwen’s open approach could become strategically important if developers rapidly build around the model.

Command: Monitor Developer Ecosystems

The strongest model is not always the model with the largest long-term influence.

Command: Track Enterprise Deployment

Real enterprise adoption will provide a better indication of practical value than launch demonstrations.

Command: Compare Chinese and American Systems

The important question is no longer simply which country has the most powerful model.

Command: Measure the Entire Stack

AI competitiveness depends on models, chips, data, infrastructure, software, talent, and distribution.

Command: Watch Semiconductor Constraints

China’s access to advanced computing remains an important factor in the long-term AI race.

Command: Watch Model Efficiency

Better algorithms can partially compensate for hardware limitations.

Command: Study Open Versus Closed AI

Open models can spread rapidly, while closed models can maintain tighter control over monetization and deployment.

Command: Measure Autonomous Productivity

A system that completes ten verified tasks overnight could be more valuable than one that achieves a spectacular benchmark score.

Command: Examine Error Costs

One serious mistake can erase the value of hundreds of successful automated actions in high-risk environments.

Command: Consider Cybersecurity

The more power AI receives, the more attractive it becomes as a target for attackers.

Command: Prepare for Agentic Malware

Attackers will eventually attempt to manipulate autonomous systems in ways similar to how they exploit human operators.

Command: Protect the AI Supply Chain

Models, plugins, tools, datasets, libraries, and external services can all become potential attack surfaces.

Command: Keep Humans Accountable

Automation should not eliminate responsibility for decisions that can seriously affect people or organizations.

Command: Rethink Productivity

The biggest economic effect may come from allowing one person to supervise several AI agents simultaneously.

Command: Watch the Cost Curve

If autonomous AI becomes cheap enough, continuous digital labor could become available to even small businesses.

Command: Expect Rapid Competition

Alibaba’s announcement will likely encourage American and Chinese competitors to emphasize their own long-running agent capabilities.

Command: Demand Independent Testing

The industry needs neutral evaluations capable of separating genuine capability from carefully engineered demonstrations.

Command: Look Beyond the Headline

“Five days of autonomous work” sounds revolutionary, but the quality of those five days is what ultimately matters.

Command: Prepare for the Agent Era

The technology is moving toward AI systems that do not simply respond to humans but actively pursue objectives on their behalf.

Command: Watch the Strategic Shift

The real AI race may ultimately be won by whoever builds the most reliable autonomous workforce rather than whoever builds the best chatbot.

What Undercode Say: The Real Battle Is No Longer About Who Has the Smartest Chatbot

Alibaba’s Qwen3.8-Max announcement is significant because it reflects where the AI industry is heading.

The most important development is not necessarily the five-day number.

It is the ambition behind the number.

Alibaba is effectively saying that AI models should be judged by how much work they can accomplish without continuous human direction.

That is a much more ambitious target than traditional conversational AI.

The comparison with Anthropic is equally important because it demonstrates that Chinese and American AI companies are converging on the same strategic objective.

Both sides are building systems designed to operate as increasingly independent digital workers.

This creates a fascinating shift in the global AI competition.

The United States may still have major advantages in advanced semiconductor access, hyperscale computing infrastructure, and frontier AI research.

China, meanwhile, has demonstrated considerable strength in rapid engineering, open-model distribution, manufacturing scale, and the ability to turn AI research into widely accessible technology.

The outcome may not be determined by one company releasing one superior model.

It could instead depend on which ecosystem can create the most capable, affordable, secure, and widely adopted AI agents.

Qwen’s open-weight strategy deserves particular attention.

If the model’s capabilities are independently confirmed and developers can deploy it without depending entirely on Alibaba’s infrastructure, its influence could extend well beyond China.

That could make Qwen strategically important even if it does not consistently outperform every proprietary American model.

The same principle applies in reverse.

American companies do not necessarily need to defeat every Chinese model on every benchmark.

They need to maintain a broader ecosystem capable of supporting frontier research, infrastructure, enterprise deployment, and continued innovation.

The emerging competition is therefore becoming multidimensional.

It involves chips.

It involves algorithms.

It involves data.

It involves research talent.

It involves open-source communities.

It involves cloud infrastructure.

It involves cybersecurity.

And increasingly, it involves autonomous digital labor.

The most important warning is that autonomy without reliability is not progress.

An AI that can operate for five days but cannot recognize when it has made a critical mistake could create more problems than it solves.

The next generation of AI evaluation must therefore reward systems that are not only capable, but also transparent, reproducible, secure, efficient, and capable of admitting failure.

That is where the real technological battle will be fought.

Qwen3.8-Max may or may not ultimately prove every claim Alibaba has made about it.

But the direction of travel is unmistakable.

AI companies are moving from assistants that answer questions toward agents that pursue objectives.

And once machines can reliably work for hours or days without continuous supervision, the economic and technological consequences could be considerably larger than anything produced by the chatbot era.

✅ Confirmed: Alibaba Reported the Long-Horizon Claims

Alibaba’s Qwen team has publicly presented Qwen3.8-Max as a model designed for coding, research, work, and long-horizon tasks, including a reported roughly five-day autonomous experiment.

⚠️ Unverified: The Performance Claims Need Independent Reproduction

The claims that the model independently worked for approximately 125 hours and outperformed hundreds of human teams are company-reported results. They should not be treated as independently verified evidence until outside researchers reproduce the experiments under comparable conditions.

✅ Confirmed Direction: Autonomous AI Is a Major Industry Trend

The broader movement toward AI agents capable of planning, coding, researching, using tools, and completing longer workflows is real. Alibaba’s announcement is part of a much larger industry shift toward increasingly autonomous AI systems.

Prediction

(+1) Autonomous AI Will Become the Next Major Competitive Metric

Over the next several years, AI companies are likely to compete less on isolated benchmark scores and more on how long their systems can complete complex, useful tasks with minimal supervision.

(+1) Open-Weight Models Will Gain Strategic Importance

If Qwen and other open models continue approaching the capabilities of proprietary systems, developers may increasingly use open-weight models as the foundation for specialized AI agents.

(+1) AI Agents Will Enter More Professional Workflows

Software development, research, data analysis, cybersecurity, administration, and technical operations are among the areas most likely to see rapid adoption of long-running AI agents.

(-1) Autonomous AI Will Also Expand the Attack Surface

As agents receive access to code, files, browsers, credentials, and business systems, attackers will have new opportunities to manipulate them. Security controls will become just as important as model intelligence.

(+1) The U.S.-China AI Competition Will Intensify

Alibaba’s Qwen strategy demonstrates that China intends to compete directly in advanced AI capabilities rather than simply follow American developments. American companies are likely to respond with increasingly autonomous systems of their own.

(-1) Marketing Claims Will Become Harder to Trust

As AI companies compete for attention, increasingly spectacular demonstrations may become common. Independent testing will be essential for determining whether headline capabilities translate into dependable real-world performance.

(+1) The Biggest Breakthrough May Be Digital Labor

If AI agents become reliable enough to perform complicated tasks independently for hours or days, they could represent a much larger economic transformation than conventional chatbots because businesses will be purchasing completed work rather than merely generated answers.

▶️ Related Video (70% Match):

🕵️‍📝Let’s dive deep and fact‑check.

🎓 Live Courses & Certifications:

Join Undercode Academy for Verified Certifications

🚀 Request a Custom Project:

Secure, high-velocity infrastructure and disruptive technological engineering. Contact our engineering team for high-tier development and proprietary systems:
[email protected]
💎 Smart Architecture | 🛡️ Secure by Design | ⭐ Trusted by Thousands

References:

Reported By: www.euronews.com
Extra Source Hub (Possible Sources for article):
https://stackoverflow.com
Wikipedia
OpenAi & Undercode AI

Image Source:

Unsplash
Undercode AI DI v2

🔐JOIN OUR CYBER WORLD [ CVE News • HackMonitor • UndercodeNews ]

💬 Whatsapp | 💬 Telegram

📢 Follow UndercodeNews & Stay Tuned:

𝕏 formerly Twitter 🐦 | @ Threads | 🔗 Linkedin | 🦋BlueSky | 🐘Mastodon | 📺Youtube