The Network Cost of an AI User
For enterprises with 600+ workers, AI does not simply add bandwidth. It changes how long flows live, how traffic moves upstream, how many sessions overlap, and which link carries the work.
A single average bandwidth number hides the variables that make AI operationally different: task size, flow duration, upload share, concurrency, transport, and where the agent runs.
AI changes the unit of network demand.
The useful planning question is not “How much bandwidth does AI use?” It is “How much traffic does this workflow generate, how long does it remain active, and where does it cross the network?”
This brief is designed for enterprises with 600 or more workers. That does not mean every employee becomes an AI user on day one. It means the access layer has to absorb adoption as it spreads across teams, floors, buildings, and branches. At that scale, aggregate concurrency—not one person’s prompt—is what turns a small workflow into a material infrastructure decision.
AI does not dominate the enterprise link today. Cisco’s report says inference traffic remains negligible beside major categories such as video in the near term. The planning signal is the combination of rapid adoption and a different traffic shape: longer connections, more upstream traffic, and more overlapping model and tool calls.
Traditional web demand is often described as short, bursty, and mostly downstream. AI inference has a different signature. In Cisco’s 2026 AI Impact on Wide Area Networks report, measured AI inference flows lasted about twice as long as non-AI web flows, while the median regular web flow rate was ten times higher. AI was not always a bigger burst. It was a smoother connection occupying the network for longer.
Direction changes too. Nine percent of measured AI inference flows were upstream-heavy, compared with roughly 0.5% of ordinary web transactions. The median downstream-to-upstream ratio narrowed from 145.39:1 for non-LLM traffic to 21.11:1 for AI traffic. That matters at branches and campuses designed around download-heavy behavior.
Transport behavior adds another planning dimension. AI services in the study used both TCP and QUIC in a nearly even flow-count split, but QUIC carried 57% of AI data volume. That can change what monitoring and inspection tools can see. The takeaway is not that every AI application has the same profile. It is that a capacity model based only on yesterday’s short, downlink-heavy web session is incomplete.
One prompt can become a network of work.
A user sees one request and one answer. The network can see a chain of model calls, retrieval, source access, tool use, and return traffic.
In Cisco’s controlled agent test, a deep-research task built with LangChain Open Deep Research analyzed 44 sources and triggered 26.8 MB of traffic across the system—450% more traffic than a comparable human-led task. Seventy percent of the new incremental traffic was AI inference.
Deep-research workflow
The essential caveat: 26.8 MB is traffic generated across this measured agent system, not a universal cost per prompt and not necessarily 26.8 MB through the user’s access switch. The agent’s location determines which model and tool calls cross each enterprise link.
That distinction prevents the most common sizing error. Multiplying total system traffic by every employee is useful for a conservative upper bound, but it can overstate the load on an access link when the agent’s fan-out happens in the cloud. It can also understate data-center or cloud-egress demand when an internally hosted agent makes many external calls.
Place AI inside the network that already exists.
A link does not carry AI in isolation. The calculator separates the AI increment from video collaboration, cloud applications, and other busy-hour traffic so you can see the whole network before and after adoption.
Size the whole link at one network boundary.
Only the 26.8 MB AI task comes from Cisco’s measured example. Duration, concurrency, link share, and the existing traffic mix are editable planning assumptions.
See the AI increment inside existing demand.
53.6% utilized after AI
10 AI tasks per user each day · 20% active at peak · one active workflow each → 120 simultaneous tasks
464.2 Mbps of modeled headroom remains. AI adds 8.6 percentage points to link utilization.
In the default “Emerging use” scenario, 20% of 600 AI-enabled workers have one active workflow at the peak. Existing busy-hour traffic uses 450 Mbps before AI. The modeled AI workload adds 85.8 Mbps, moving the whole link from 45% to 53.6% utilization.
Move the intensity slider one step at a time to test the range. Its five named milestones describe recognizable workload patterns; every position between them smoothly interpolates task frequency, peak-active share, and parallel workflows. “Minimal” assumes six AI tasks per worker each day and only 5% active at peak. “Intensive” is a stress-test ceiling: 288 tasks per worker each day, every worker active at peak, and three parallel AI workflows or applications per active worker. These are transparent planning profiles—not an adoption forecast.
The existing video, cloud, and other traffic defaults are not vendor forecasts or per-employee constants. Replace them with observed busy-hour throughput at the same switch uplink, WAN circuit, or security boundary being modeled. Keep the categories separate so the AI increment remains visible instead of disappearing inside one total.
Below the peak comparison, the ledger still shows AI volume: ten tasks produce 268 MB per AI-enabled worker per workday, and 600 adopted users produce 160.8 GB per day across the full system. Peak Mbps and daily GB answer different questions, so the calculator keeps both visible without pretending they are interchangeable.
AI volume compounds across adopted users.
These rows isolate AI traffic and reuse the calculator’s tasks-per-day and MB-per-task values. Monthly totals assume 20 workdays.
| AI-enabled users | Per user / day | Site / day | Site / 20 days |
|---|---|---|---|
| 600 | 268 MB | 160.8 GB | 3.22 TB |
| 1,500 | 268 MB | 402 GB | 8.04 TB |
| 5,000 | 268 MB | 1.34 TB | 26.8 TB |
Where the bytes land matters more than the total.
The same agent task can be inexpensive at the office edge and expensive at enterprise egress—or the reverse—depending on where orchestration, models, and tools run.
If an agent runs in a cloud service, the office access link may carry only the user’s request and the final streamed response. Most of the agent-to-model and agent-to-tool traffic can remain inside the provider’s cloud. If the agent runs in an enterprise data center, those fan-out calls may cross a WAN, firewall, proxy, private interconnect, or internet egress point instead.
Measure at the boundary you intend to protect or expand.
Placement also changes troubleshooting. A slow answer from an internal AI assistant can still be caused by an oversubscribed access-switch uplink even when the model is healthy. Conversely, an office link can look healthy while an enterprise egress point, security stack, or cloud interconnect is carrying the agent’s hidden fan-out. Total bytes describe the workload. Link-local measurements identify the bottleneck.
The quiet risk is concurrent state.
A lower-rate flow is not necessarily a lighter operational burden when it remains open longer and overlaps with hundreds of other tasks.
Longer-lived AI connections change more than utilization. They can increase the number of simultaneous sessions held by firewalls, proxies, network address translation, load balancers, and monitoring systems. Agentic workflows add fan-out: one user task can create multiple model and tool conversations before the answer returns.
Different flow shape, different operating pressure.
That is why average daily volume and peak Mbps belong in the same model but answer different questions. Daily volume informs transit, egress, and cost. Peak Mbps informs immediate link pressure. Concurrent flows, duration, and transport mix inform the stateful infrastructure in between. None should be inferred from only one of the others.
Latency needs similar care. Cisco notes that today’s model inference time can dominate the network—for example, tens of milliseconds of network latency inside a response that takes seconds to generate. As inference hardware becomes faster, the network becomes a larger share of the experience. The practical response is not to blame or absolve the network in advance. It is to measure model time and path time together.
Turn the workload change into an access-layer decision.
AI strengthens the case for a faster, more efficient access layer—but the upgrade trigger is measured aggregate demand, not the word “AI” on an application roadmap.
The traffic profile is changing in several directions at once: longer-lived sessions, more upstream context, more concurrent flows, and agent fan-out across models and tools. Across a 600+ worker enterprise, those characteristics join video, cloud applications, and device growth on the same shared wireless and switching infrastructure. The relevant unit becomes the combined busy-hour load across a floor, building, or site—not a single user’s average.
Cisco’s Wi-Fi 7 design guide describes the technologies that expand that envelope: 320 MHz channels can double potential throughput versus 160 MHz, 4096-QAM carries 20% more data per symbol under suitable radio conditions, and Multi-Link Operation can use multiple bands to improve throughput, latency, and reliability. These are shared-radio capabilities—not guaranteed application speeds for every client.
Busy-hour channel utilization, retries, client density, roaming, latency, and 6 GHz-capable client share.
Wi-Fi 7 design using 6 GHz, MLO, wider channels where appropriate, and a channel plan matched to the environment.
AP switch-port throughput, burst behavior, queue drops, negotiated speed, and power mode.
2.5, 5, or 10 GbE multigig ports so the wired handoff does not cap aggregate wireless capacity.
Access-switch uplink utilization, oversubscription, buffer pressure, and simultaneous AI task volume.
Higher-capacity switch uplinks or fabric sized for the combined busy-hour load—not the rating of one AP.
Link-local AI traffic share, upstream/downstream balance, flow duration, loss, latency, and inspection load.
WAN, egress, or security capacity at the constrained boundary. A wireless refresh will not repair a bottleneck elsewhere.
The wired side has to keep up with the radios. Cisco’s current Wi-Fi 7 portfolio illustrates the range: the CW9172 uses a 2.5 GbE multigig uplink, while the higher-capacity CW9176 supports 2.5, 5, and 10 GbE. Cisco’s design guide notes that capable Wi-Fi 7 access points can exceed a 1 GbE handoff under aggregate load. That is the practical case for multigig access switching: prevent the AP’s switch port from becoming the ceiling after investing in more wireless capacity.
This does not mean every branch needs 10 GbE to every AP today. Client support, channel availability, RF design, cabling, Power over Ethernet, switch uplinks, WAN capacity, and actual workload concurrency still determine the result. Upgrade the layer the evidence identifies, then rerun the same representative AI tasks to prove the user experience changed.
Choose a representative office or branch from the 600+ worker environment. Measure three real AI workflows during the busy hour, trace every boundary they cross, and compare the result with wireless airtime, AP ports, switch uplinks, security inspection, and WAN headroom.
Build the baseline before adoption hides it.
The goal is not one universal “AI bandwidth” number. It is a repeatable measurement method for each important workflow and network boundary.
Separate simple chat, retrieval, coding, document analysis, image generation, and agentic research. Each has a different task size, duration, and fan-out pattern.
Capture bytes, duration, upstream/downstream balance, flows, transport, and user-perceived time. Repeat enough times to understand median and high-end behavior.
Measure access, WAN, enterprise egress, data-center, and cloud paths separately. The total workflow and the load on one link are not the same quantity.
Use task starts by minute and active duration—not just licensed users—to establish a peak planning envelope. Recheck after rollout changes user behavior.
Connect flow and link data with model runtime, tool calls, errors, and experience. ThousandEyes network tests, for example, expose loss, latency, jitter, throughput, and hop-by-hop path evidence that can separate link pressure from model delay.
The result should be a living workload profile: MB per task, active seconds per task, upload ratio, concurrent flows, QUIC/TCP mix, path placement, peak starts per minute, and user experience. Rerun it when the model, agent design, tool set, data source, or hosting pattern changes.
Start with one real task. Trace the traffic it creates. Identify the boundary that carries it. Then scale the math with observed usage instead of a generic assumption. That is how an AI roadmap becomes a network plan.
Source Trail View the Evidence and Calculation Method Measured findings are kept separate from illustrative calculations and author interpretation.
Cisco: AI Impact on Wide Area Networks
Primary source for the live traffic comparison, near-term scale, flow duration and rate, upstream-heavy flow share, down/up ratios, TCP/QUIC mix, and the controlled deep-research test.
ThousandEyes Network Tests
Primary product documentation for the network measurements referenced in the operational playbook: loss, latency, jitter, throughput, path trace, hop IPs, and per-hop latency.
ThousandEyes Path Visualization
Primary product documentation for visualizing end-to-end paths and the network nodes between an agent and target.
Cisco Wi-Fi 7 Design Guide
Primary source for 4096-QAM, 320 MHz channels, Multi-Link Operation, access-layer design considerations, and the relationship between high-capacity Wi-Fi 7 radios and switch-port speed.
Cisco Wi-Fi 7 Access Points
Primary product data for the concrete multigig examples: a 2.5 GbE uplink on the CW9172 and a 2.5/5/10 GbE multigig uplink on the CW9176.
Calculator Method
Per-user volume equals tasks per day × MB per task. Active-task Mbps equals MB per task × 8 ÷ task seconds × link share. Peak link load equals active-task Mbps × simultaneous tasks. Monthly totals use 20 workdays and decimal units.
Interpretation Boundary
Cisco’s 26.8 MB result describes one controlled deep-research workflow across the measured system. This article uses it as an editable example, not a universal AI-user average or product sizing recommendation.
React or leave a comment.
What are you measuring as AI traffic moves from experiments into daily work?