Agentic AI in China: An Enterprise Field Guide
Evaluate enterprise AI agents through tool permissions, approval boundaries, recovery from failed tasks and the real cost of human supervision.

An enterprise agent should be judged by the work it can complete within defined permissions. Distinguish suggestions from actions, reversible actions from commitments, and unattended execution from work that relies on hidden human assistance. Ask for the unsuccessful runs as well as the successful demonstration.
What to evaluate
Workflow depth
Trace goals, context, decisions, tool calls, approvals and final outcomes.
Control boundary
Inspect identity, permissions, budgets, validation, logs and reversible actions.
Reliability
Measure completion, intervention, retries, latency, cost and failure recovery.
Operating ownership
Identify who maintains prompts, tools, knowledge, evaluation and incidents.
Worked evaluation exercise
The following is an illustrative assessment exercise, not a reported customer result.
For a purchase-request assistant, begin in a sandbox with synthetic suppliers and a spending limit. Require human approval before creating any external commitment. Introduce a missing delivery date and a conflicting supplier document, then examine whether the agent asks for clarification or invents an answer. Count retries, review time and abandoned tasks in the cost per completed request. Those measures describe the workflow more clearly than the number of tools the agent can call.
Questions for the operating team
- What can the agent change without approval?
- How often does a human intervene?
- Can every external action be reconstructed and reversed?
Warning signs
- One high-privilege service account.
- Successful demos with no failure taxonomy.
- Human review reduced to a default confirmation.
Planning the visit
Use these questions to scope an industry-focused China AI program. Agree which records and operating workflows can be examined before confirming meetings; host participation and access require confirmation. Our research method explains how we structure the brief and follow-up.
References and scope
These references provide policy or risk-management context. They do not independently verify a host's performance or the illustrative exercise above.