Field Brief AI Capacity Planning July 2026

The Network Cost of an AI User

For enterprises with 600+ workers, AI does not simply add bandwidth. It changes how long flows live, how traffic moves upstream, how many sessions overlap, and which link carries the work.

In brief Capacity planning for AI starts with one task, one path, and one honest answer about which bytes cross the link you are sizing.

A single average bandwidth number hides the variables that make AI operationally different: task size, flow duration, upload share, concurrency, transport, and where the agent runs.

AI changes the unit of network demand.

The useful planning question is not “How much bandwidth does AI use?” It is “How much traffic does this workflow generate, how long does it remain active, and where does it cross the network?”

This brief is designed for enterprises with 600 or more workers. That does not mean every employee becomes an AI user on day one. It means the access layer has to absorb adoption as it spreads across teams, floors, buildings, and branches. At that scale, aggregate concurrency—not one person’s prompt—is what turns a small workflow into a material infrastructure decision.

AI does not dominate the enterprise link today. Cisco’s report says inference traffic remains negligible beside major categories such as video in the near term. The planning signal is the combination of rapid adoption and a different traffic shape: longer connections, more upstream traffic, and more overlapping model and tool calls.

Traditional web demand is often described as short, bursty, and mostly downstream. AI inference has a different signature. In Cisco’s 2026 AI Impact on Wide Area Networks report, measured AI inference flows lasted about twice as long as non-AI web flows, while the median regular web flow rate was ten times higher. AI was not always a bigger burst. It was a smoother connection occupying the network for longer.

Direction changes too. Nine percent of measured AI inference flows were upstream-heavy, compared with roughly 0.5% of ordinary web transactions. The median downstream-to-upstream ratio narrowed from 145.39:1 for non-LLM traffic to 21.11:1 for AI traffic. That matters at branches and campuses designed around download-heavy behavior.

Observed traffic fingerprint Regular web AI inference
Flow duration How long the connection stays active
normalized baseline
≈2× approximately twice as long
Median flow rate Throughput while the flow is active
10× higher median rate
lower and smoother
Upstream-heavy flows Flows sending more than they receive
≈0.5% ordinary web transactions
9% measured AI inference flows
Down / up ratio Median traffic direction balance
145:1 strongly downstream
21:1 meaningfully more symmetric
Cisco’s measurements compare live AI inference traffic with non-AI web traffic. The multiples are normalized to make the behavioral difference readable; they are not universal sizing constants.

Transport behavior adds another planning dimension. AI services in the study used both TCP and QUIC in a nearly even flow-count split, but QUIC carried 57% of AI data volume. That can change what monitoring and inspection tools can see. The takeaway is not that every AI application has the same profile. It is that a capacity model based only on yesterday’s short, downlink-heavy web session is incomplete.

One prompt can become a network of work.

A user sees one request and one answer. The network can see a chain of model calls, retrieval, source access, tool use, and return traffic.

In Cisco’s controlled agent test, a deep-research task built with LangChain Open Deep Research analyzed 44 sources and triggered 26.8 MB of traffic across the system—450% more traffic than a comparable human-led task. Seventy percent of the new incremental traffic was AI inference.

Measured task anatomy One visible request. Multiple network conversations.
Cisco controlled test
Deep-research workflow
01 · User Research request One action starts the workflow.
02 · Agent Plans and orchestrates 26.8 MB total traffic triggered across the system
03 · Model Inference calls 70% of incremental traffic
04 · Tools Retrieval and sources 44 sources analyzed
05 · User Final response The small visible answer is not the full workload.

The essential caveat: 26.8 MB is traffic generated across this measured agent system, not a universal cost per prompt and not necessarily 26.8 MB through the user’s access switch. The agent’s location determines which model and tool calls cross each enterprise link.

Measured values are from Cisco’s 2026 report. The route is a conceptual decomposition of the tested workflow, not a packet-level reconstruction of Cisco’s lab.

That distinction prevents the most common sizing error. Multiplying total system traffic by every employee is useful for a conservative upper bound, but it can overstate the load on an access link when the agent’s fan-out happens in the cloud. It can also understate data-center or cloud-egress demand when an internally hosted agent makes many external calls.

Place AI inside the network that already exists.

A link does not carry AI in isolation. The calculator separates the AI increment from video collaboration, cloud applications, and other busy-hour traffic so you can see the whole network before and after adoption.

Interactive capacity ledger

Size the whole link at one network boundary.

Only the 26.8 MB AI task comes from Cisco’s measured example. Duration, concurrency, link share, and the existing traffic mix are editable planning assumptions.

Existing busy-hour traffic Illustrative Mbps at this boundary—replace with measured telemetry
Whole-link busy hour

See the AI increment inside existing demand.

53.6% utilized after AI

AI integration intensity Emerging use

10 AI tasks per user each day · 20% active at peak · one active workflow each → 120 simultaneous tasks

Move one step at a time; milestone labels mark the modeled workload pattern.
Before AI 450 Mbps 45% of link capacity
AI adds +85.8 Mbps 16% of combined peak
After AI 535.8 Mbps 53.6% of link capacity

464.2 Mbps of modeled headroom remains. AI adds 8.6 percentage points to link utilization.

Per user / day 268 MB tasks × traffic per task
Site / workday 160.8 GB before the link-share adjustment
Site / 20 workdays 3.22 TB illustrative monthly total
One active task on link 0.71 Mbps average during the entered duration
AI peak = 26.8 MB × 8 ÷ 300 sec × 120 simultaneous × 100% link share Whole link = 180 video + 160 cloud + 110 other + 85.8 AI = 535.8 Mbps

In the default “Emerging use” scenario, 20% of 600 AI-enabled workers have one active workflow at the peak. Existing busy-hour traffic uses 450 Mbps before AI. The modeled AI workload adds 85.8 Mbps, moving the whole link from 45% to 53.6% utilization.

Move the intensity slider one step at a time to test the range. Its five named milestones describe recognizable workload patterns; every position between them smoothly interpolates task frequency, peak-active share, and parallel workflows. “Minimal” assumes six AI tasks per worker each day and only 5% active at peak. “Intensive” is a stress-test ceiling: 288 tasks per worker each day, every worker active at peak, and three parallel AI workflows or applications per active worker. These are transparent planning profiles—not an adoption forecast.

The existing video, cloud, and other traffic defaults are not vendor forecasts or per-employee constants. Replace them with observed busy-hour throughput at the same switch uplink, WAN circuit, or security boundary being modeled. Keep the categories separate so the AI increment remains visible instead of disappearing inside one total.

Below the peak comparison, the ledger still shows AI volume: ten tasks produce 268 MB per AI-enabled worker per workday, and 600 adopted users produce 160.8 GB per day across the full system. Peak Mbps and daily GB answer different questions, so the calculator keeps both visible without pretending they are interchangeable.

Scale effect

AI volume compounds across adopted users.

These rows isolate AI traffic and reuse the calculator’s tasks-per-day and MB-per-task values. Monthly totals assume 20 workdays.

AI-enabled users Per user / day Site / day Site / 20 days
600 268 MB 160.8 GB 3.22 TB
1,500 268 MB 402 GB 8.04 TB
5,000 268 MB 1.34 TB 26.8 TB

Where the bytes land matters more than the total.

The same agent task can be inexpensive at the office edge and expensive at enterprise egress—or the reverse—depending on where orchestration, models, and tools run.

If an agent runs in a cloud service, the office access link may carry only the user’s request and the final streamed response. Most of the agent-to-model and agent-to-tool traffic can remain inside the provider’s cloud. If the agent runs in an enterprise data center, those fan-out calls may cross a WAN, firewall, proxy, private interconnect, or internet egress point instead.

Traffic placement

Measure at the boundary you intend to protect or expand.

01 · User edge
Device → access → WAN Always sees the user request and returned answer. It sees fan-out only when the agent is on the far side of this boundary.
Access-link measurement
02 · Enterprise boundary
Agent → security → egress Can carry repeated model, retrieval, and tool calls when the agent is hosted inside the enterprise.
Egress measurement
03 · Service fabric
Models, tools, and sources Cloud-hosted fan-out may never return through the office link even though it is part of the total system workload.
The calculator’s “traffic crossing this link” input exists for this reason. Set it from observation where possible; use 100% only as a conservative bound when architecture details are not yet known.

Placement also changes troubleshooting. A slow answer from an internal AI assistant can still be caused by an oversubscribed access-switch uplink even when the model is healthy. Conversely, an office link can look healthy while an enterprise egress point, security stack, or cloud interconnect is carrying the agent’s hidden fan-out. Total bytes describe the workload. Link-local measurements identify the bottleneck.

The quiet risk is concurrent state.

A lower-rate flow is not necessarily a lighter operational burden when it remains open longer and overlaps with hundreds of other tasks.

Longer-lived AI connections change more than utilization. They can increase the number of simultaneous sessions held by firewalls, proxies, network address translation, load balancers, and monitoring systems. Agentic workflows add fan-out: one user task can create multiple model and tool conversations before the answer returns.

Connection occupancy

Different flow shape, different operating pressure.

Regular web short, higher-rate bursts
AI inference lower-rate, longer-lived flow
This is a qualitative timeline based on Cisco’s observed duration and median-rate relationship. It illustrates occupancy, not an exact byte-for-byte trace.

That is why average daily volume and peak Mbps belong in the same model but answer different questions. Daily volume informs transit, egress, and cost. Peak Mbps informs immediate link pressure. Concurrent flows, duration, and transport mix inform the stateful infrastructure in between. None should be inferred from only one of the others.

Latency needs similar care. Cisco notes that today’s model inference time can dominate the network—for example, tens of milliseconds of network latency inside a response that takes seconds to generate. As inference hardware becomes faster, the network becomes a larger share of the experience. The practical response is not to blame or absolve the network in advance. It is to measure model time and path time together.

Turn the workload change into an access-layer decision.

AI strengthens the case for a faster, more efficient access layer—but the upgrade trigger is measured aggregate demand, not the word “AI” on an application roadmap.

The traffic profile is changing in several directions at once: longer-lived sessions, more upstream context, more concurrent flows, and agent fan-out across models and tools. Across a 600+ worker enterprise, those characteristics join video, cloud applications, and device growth on the same shared wireless and switching infrastructure. The relevant unit becomes the combined busy-hour load across a floor, building, or site—not a single user’s average.

Cisco’s Wi-Fi 7 design guide describes the technologies that expand that envelope: 320 MHz channels can double potential throughput versus 160 MHz, 4096-QAM carries 20% more data per symbol under suitable radio conditions, and Multi-Link Operation can use multiple bands to improve throughput, latency, and reliability. These are shared-radio capabilities—not guaranteed application speeds for every client.

Observed constraint Evidence to collect Action to evaluate
Wireless airtime or contention

Busy-hour channel utilization, retries, client density, roaming, latency, and 6 GHz-capable client share.

Wi-Fi 7 design using 6 GHz, MLO, wider channels where appropriate, and a channel plan matched to the environment.

AP wired uplink near 1 Gbps

AP switch-port throughput, burst behavior, queue drops, negotiated speed, and power mode.

2.5, 5, or 10 GbE multigig ports so the wired handoff does not cap aggregate wireless capacity.

Many APs converge on one switch

Access-switch uplink utilization, oversubscription, buffer pressure, and simultaneous AI task volume.

Higher-capacity switch uplinks or fabric sized for the combined busy-hour load—not the rating of one AP.

WAN or security boundary is slow

Link-local AI traffic share, upstream/downstream balance, flow duration, loss, latency, and inspection load.

WAN, egress, or security capacity at the constrained boundary. A wireless refresh will not repair a bottleneck elsewhere.

The wired side has to keep up with the radios. Cisco’s current Wi-Fi 7 portfolio illustrates the range: the CW9172 uses a 2.5 GbE multigig uplink, while the higher-capacity CW9176 supports 2.5, 5, and 10 GbE. Cisco’s design guide notes that capable Wi-Fi 7 access points can exceed a 1 GbE handoff under aggregate load. That is the practical case for multigig access switching: prevent the AP’s switch port from becoming the ceiling after investing in more wireless capacity.

This does not mean every branch needs 10 GbE to every AP today. Client support, channel availability, RF design, cabling, Power over Ethernet, switch uplinks, WAN capacity, and actual workload concurrency still determine the result. Upgrade the layer the evidence identifies, then rerun the same representative AI tasks to prove the user experience changed.

Call to action Build one AI-ready access baseline this quarter.

Choose a representative office or branch from the 600+ worker environment. Measure three real AI workflows during the busy hour, trace every boundary they cross, and compare the result with wireless airtime, AP ports, switch uplinks, security inspection, and WAN headroom.

Build the baseline before adoption hides it.

The goal is not one universal “AI bandwidth” number. It is a repeatable measurement method for each important workflow and network boundary.

01 Define representative work.

Separate simple chat, retrieval, coding, document analysis, image generation, and agentic research. Each has a different task size, duration, and fan-out pattern.

02 Measure one completed task.

Capture bytes, duration, upstream/downstream balance, flows, transport, and user-perceived time. Repeat enough times to understand median and high-end behavior.

03 Observe the right boundary.

Measure access, WAN, enterprise egress, data-center, and cloud paths separately. The total workflow and the load on one link are not the same quantity.

04 Model actual concurrency.

Use task starts by minute and active duration—not just licensed users—to establish a peak planning envelope. Recheck after rollout changes user behavior.

05 Correlate path and workload.

Connect flow and link data with model runtime, tool calls, errors, and experience. ThousandEyes network tests, for example, expose loss, latency, jitter, throughput, and hop-by-hop path evidence that can separate link pressure from model delay.

The result should be a living workload profile: MB per task, active seconds per task, upload ratio, concurrent flows, QUIC/TCP mix, path placement, peak starts per minute, and user experience. Rerun it when the model, agent design, tool set, data source, or hosting pattern changes.

AI capacity is not a single forecast. It is a measured relationship between work, time, concurrency, and place.

Start with one real task. Trace the traffic it creates. Identify the boundary that carries it. Then scale the math with observed usage instead of a generic assumption. That is how an AI roadmap becomes a network plan.

Source Trail View the Evidence and Calculation Method Measured findings are kept separate from illustrative calculations and author interpretation.

Cisco: AI Impact on Wide Area Networks

Primary source for the live traffic comparison, near-term scale, flow duration and rate, upstream-heavy flow share, down/up ratios, TCP/QUIC mix, and the controlled deep-research test.

ThousandEyes Network Tests

Primary product documentation for the network measurements referenced in the operational playbook: loss, latency, jitter, throughput, path trace, hop IPs, and per-hop latency.

ThousandEyes Path Visualization

Primary product documentation for visualizing end-to-end paths and the network nodes between an agent and target.

Cisco Wi-Fi 7 Design Guide

Primary source for 4096-QAM, 320 MHz channels, Multi-Link Operation, access-layer design considerations, and the relationship between high-capacity Wi-Fi 7 radios and switch-port speed.

Cisco Wi-Fi 7 Access Points

Primary product data for the concrete multigig examples: a 2.5 GbE uplink on the CW9172 and a 2.5/5/10 GbE multigig uplink on the CW9176.

Calculator Method

Per-user volume equals tasks per day × MB per task. Active-task Mbps equals MB per task × 8 ÷ task seconds × link share. Peak link load equals active-task Mbps × simultaneous tasks. Monthly totals use 20 workdays and decimal units.

Transparent arithmetic · no forecast model

Interpretation Boundary

Cisco’s 26.8 MB result describes one controlled deep-research workflow across the measured system. This article uses it as an editable example, not a universal AI-user average or product sizing recommendation.

Author analysis
Join the thread

React or leave a comment.

What are you measuring as AI traffic moves from experiments into daily work?