China Edge AI Visits: Separate Latency from Throughput
Latency, throughput and end-to-end response time answer different questions. An edge AI visit should pin down the measurement boundary before comparing hardware.

Direct answer. Ask where the timer starts and stops, how many tasks run concurrently and what happened to slow or failed requests. A quoted inference time alone does not describe the response that an operator or machine receives.
Draw the measurement boundary
On a China edge AI company visit, follow one input through capture, preprocessing, model inference and delivery of the result. If the demonstration operates machinery, include only the host-authorized explanation of acknowledgement; visitors should never initiate control actions. The hardware on the table may accelerate only one segment of a longer workflow.
Ask whether the figures describe a warm system, the first request after startup or sustained load. Note model version, precision setting, input size and device configuration. These are comparison conditions, not minor implementation details. A different model or workload can invalidate an apparently neat hardware ranking.
Separate the metrics
| Metric | What it answers | Follow-up question |
|---|---|---|
| Inference latency | Time inside the measured model stage | Were preprocessing and queueing excluded? |
| Throughput | Completed work per unit time | At what concurrency and error rate? |
| End-to-end delay | Time until the usable result arrives | Were capture and network delivery included? |
| Tail behavior | Slower outcomes within the workload | Were timeouts counted or discarded? |
Prefer a workload description and a latency distribution to one headline average. If only an average is available, record that limitation. Do not manufacture percentile values from a small demonstration or assume an excluded timeout completed successfully.
The request for documented testing conditions is aligned with NIST AI RMF Core, notably MEASURE 2.1 and 2.3. The benchmark worksheet is our own visit aid, not a NIST performance standard or certification.
Turn a hardware impression into a pilot question
Ask the host to quote the same workload on the proposed production configuration, including thermal conditions, available power, network interruptions and maintenance constraints. An attractive desktop demonstration does not establish the result inside your enclosure or site.
A useful visit conclusion names the measured stage and leaves the unmeasured stages open. Refer to the company due-diligence checklist, then brief an executive immersion with your engineering team's actual timing requirements.