AI & CX TechnologyAI in CXAI InfrastructureArtificial IntelligenceContact Center InnovationCustomer Experience (CX)Expert OpinionsLeadershipLeadership InsightsTechnology

AI Economics in the Contact Center: What a Four-H200 Test Revealed About Cost, Capacity and Risk

An operator case study from The Call Center Doctors on renting GPUs, running DeepSeek-V4.1-Flash, and why contact centers should measure AI by business outcomes rather than token prices.

By Jason Shouldice, founder and CEO, The Call Center Doctors

At 12:00 on September 27, 2026, a four-GPU server was still costing money even though it had stopped doing useful work.

The machine had four NVIDIA H200 GPUs. The rental meter was running at $9.186 an hour. The model serving stack had entered a CUDA graph capture process and stopped responding without producing a useful error.

The instinct was to restart it.

Instead, we decided to let it run and investigate what had gone wrong.

That decision eventually produced a three-hour experiment into a question that contact-center technology leaders increasingly face:

When does running an open AI model on rented infrastructure make economic and operational sense compared with using a frontier model through an API or subscription?

The experiment was not conducted on customer calls or customer data.

It used our internal software-engineering workload, where AI coding agents help us build the contact-center software we use for dialer operations, inbound answering, call grading and AI-assisted workflows.

That distinction matters.

The measurements in this case study describe our engineering workload. They do not constitute a controlled benchmark of voice AI, contact-center QA or customer-facing model performance.

But the operational questions travel.

If a contact center is considering self-hosted or privately deployed AI for voice agents, quality assurance, call summarisation or other workloads, the economics cannot be reduced to the price per million tokens.

GPU utilisation, workload composition, concurrency, infrastructure expertise, security controls, model capability and—most importantly—the business unit being produced all affect the answer.

Why a Contact Center Was Renting GPUs

We are a US contact center.

We set appointments on outbound campaigns, answer inbound lines and run quality assurance on calls. And, we also build our own contact-center software in-house, including dialer operations, call grading and an AI assistant.

About ten full-time AI developers work on that software, mostly by directing AI coding agents. Claude Code is our primary development environment.

That means AI model consumption is an operating cost.

Between September 1 and September 27, 2026, the agents on one development server made 2,031,901 model calls, while the team merged 5,610 code changes.

At that scale, the question becomes practical rather than theoretical:

Should we run an open model on GPUs we rent, or should we continue paying a frontier-model provider?

We chose a narrow experiment.

We would rent GPUs, run DeepSeek-V4.1-Flash, connect it to our existing agent workflow and measure what happened.

No customer calls were used.

No customer data was used.

The purpose was to understand the economics and operational burden of the infrastructure.

The Configuration: Four H200 GPUs and DeepSeek-V4.1-Flash

Eight-card H200 spot instances were unavailable that morning. A four-card machine became available, and we took it.

The machine came from Verda and contained:

  • 4 × NVIDIA H200 GPUs
  • 141 GB GPU memory per card
  • $9.186 per hour
  • Approximately $220 for 24 hours

The model was DeepSeek-V4.1-Flash, as published on Hugging Face.

The model files totalled approximately 511 GB. We counted 763 billion parameters in the files, with approximately 552 billion in the main model.

We served the model using vLLM nightly 0.30.1rc1, because the stable release available at the time could not load the V4.1 architecture.

Tensor parallelism was set to four, meaning one model instance was distributed across all four GPUs.

The model’s mixture-of-experts weights were supplied in MXFP4. vLLM converted them to 8-bit during loading. The Engram memory module operated in CPU RAM.

The serving configuration also included:

  • DSpark speculative decoding
  • Five speculative tokens
  • Adaptive verification disabled
  • 90% GPU memory utilisation
  • 8,192-token text chunks
  • Prefix caching enabled
  • Maximum 128 concurrent requests
  • Full 1,048,576-token context
  • Claude Code connecting through a local Anthropic-compatible gateway

Our baseline was the model environment developers use every day: Claude Opus 5.5 through our existing subscriptions.

This was therefore an infrastructure experiment, not a laboratory benchmark.

Three Failed Starts Before the Model Worked

The first start began at 11:21.

At 11:27, the progress indicator reported that all 48 files had loaded.

It was wrong.

The weights actually finished loading at 11:35 because the expert weights were converted to 8-bit on the CPU during startup.

CUDA graph capture then consumed approximately 42 GB per GPU.

At 11:43, the process failed with:

“No available memory for the cache blocks.”

Forty-one minutes after renting the machine, we had no usable model.

The second start froze around noon.

Research into the configuration pointed toward adaptive verification as one possible cause.

A third start, based on a publicly available H200 benchmark configuration, failed because a memory setting intended for H100 hardware interfered with the fast inter-GPU communication path.

The fourth start came up at 13:05.

The model answered its first question in 1.7 seconds, made a correct tool call and, through Claude Code, produced a working fib.py.

It had taken more than two hours from rental to the first successful response.

That is an important part of the economics.

Infrastructure expertise is itself a cost.

An API customer does not normally spend an afternoon diagnosing GPU memory configuration, serving software compatibility, speculative-decoding settings or tensor-parallel behaviour.

A self-hosted deployment does.

Finding the Capacity Ceiling

Once the configuration was stable, we pushed it.

A counter reached 7,389 tokens per second in bursts at 256 concurrent streams.

Minutes later, 256 streams requesting 1,500-token responses crashed the box.

We had found the ceiling by exceeding it.

We then capped concurrency at 128 streams.

Start five became stable at 13:42.

We ran four one-minute saturation tests with zero errors.

The measured throughput was:

MeasureResult
New text16,621 tokens/sec
Cached text521,027 tokens/sec
Written output5,281 tokens/sec
Mixed workload11,196 new / 131,904 cached / 521 written tokens/sec

That distinction between new and cached text is particularly important.

Not all tokens cost the same amount of infrastructure time.

Our coding agents repeatedly reread large contexts. In our September workload, 96.3% of the input tokens were cached context.

But each call also introduced approximately 7,800 new tokens.

In our measurements, a new token consumed roughly 30 times as much box time as a cached token.

Using our capacity model, checked against the live mixed-workload test to within approximately 3%, the real workload reached a ceiling of around 213 written tokens per second, equivalent to roughly 20 billion tokens per day under the model’s assumptions.

That number is useful only in context.

It is not a general DeepSeek capacity claim.

It describes what our four-GPU configuration could support under our workload assumptions.

The Economics Look Different When Utilisation Changes

We then compared the infrastructure cost with the equivalent DeepSeek API consumption.

A full day of our workload at DeepSeek’s own API pricing would have been approximately:

  • $184 at off-peak rates
  • $223 when accounting for actual peak-hour pricing

The GPU rental was approximately $220 per day.

At 100% utilisation, the rented box therefore roughly matched the API economics for our workload.

But contact-center infrastructure is rarely 100% busy all day.

An idle GPU still costs money.

That is similar to a staffed seat with no calls in the queue: the capacity has been purchased whether or not it is being consumed.

Our average September day would have kept the box approximately 57–71% utilised, depending on which work was included.

At the lower end of that range, the rented infrastructure cost roughly twice the equivalent API consumption.

Our busiest day, September 23, would have required approximately two and a half such boxes.

This is why a simple per-token comparison can be misleading.

The infrastructure question is not only:

How much does a token cost?

It is also:

How consistently can I keep the infrastructure productive?

The Latency Comparison Was Not Apples to Apples

We also measured response latency against our Claude environment.

But this was not a controlled benchmark.

The two environments differed in workload, context size, concurrency, tools and infrastructure.

On the Claude side, we measured 26,680 real calls from our Claude Code sessions doing everyday engineering work with tools enabled.

Median context was approximately 542,165 tokens.

On the DeepSeek side, we measured 853 calls from read-only code-review agents, with 48–64 agents sharing the same four-GPU machine.

Median context was approximately 24,456 tokens.

The measured figures were:

MeasureClaude Opus 5.5DeepSeek on our H200 box
Median first response3.5 sec12.8 sec
90th percentile9.6 sec56.6 sec
Approx. output rate per call85 tok/sec8 tok/sec

Earlier, with only 16 concurrent streams, DeepSeek had reached approximately 105 output tokens per second per call.

The difference is partly explained by concurrency and infrastructure.

Claude requests were distributed across a large commercial provider’s infrastructure.

Our DeepSeek workload was concentrated on one four-GPU machine.

Therefore, these numbers should be read as operator experience under different configurations, not as a model-quality benchmark.

The experiment itself made the limitation obvious.

As we noted during testing, it was almost an apples-to-oranges comparison.

What the Published Benchmarks Actually Say

We did not independently measure model quality.

The frequently cited 66 versus 31 comparison in our original material is vendor-reported.

On Terminal-Bench 4.0 at maximum effort:

  • Anthropic reports 66.4 for Claude Opus 5.5.
  • DeepSeek’s model card reports 31.2 for DeepSeek-V4.1-Flash.

Those numbers should not be treated as a controlled head-to-head measurement by us.

Each vendor ran its own evaluation setup.

DeepSeek’s model card also reports 74.2 for V4.1-Flash on DeepSWE v1.1 and lists Opus 5 at 74.0 on that evaluation.

The broader point for an enterprise buyer is therefore not that one number settles model selection.

It is that benchmark results can vary substantially by task and evaluation harness.

A contact center considering an AI model should test the workload it actually intends to run.

The Security Problem Was Ours

The model was not the only engineering challenge.

Before an agent could write real code, it needed a restricted execution environment.

We built that sandbox ourselves as part of the orchestration layer that runs multiple agents on a shared server.

The environment used a dedicated Linux user, restricted network access and no access to production credentials.

Two independent reviewers tested the sandbox through six rounds of fixes.

They found four ways around the restrictions.

One vulnerability involved a launcher trusting a settings file in a shared temporary directory that another user on the server could modify.

The important point is that these were vulnerabilities in our orchestration code, not in the model.

We found them during review, before any code-writing agent was permitted to run.

Because of those vulnerabilities—and because we had a separate rule that agent-written code must never execute as a real person—the DeepSeek code-writing agents never went live.

Only read-only DeepSeek reviewers ran.

DeepSeek therefore shipped zero lines of code into our product during this experiment.

That is another lesson for contact-center AI deployments.

Running an open model yourself does not eliminate security responsibility.

It changes where that responsibility sits.

Whether the workload involves code, call transcripts, recordings, customer information or internal systems, the organisation operating the AI stack also owns the surrounding controls.

From Cost Per Token to Cost Per Business Outcome

During the test, one thought kept coming back:

“I mean even if it’s half the price for coding.”

That is the wrong stopping point.

For a contact center, the relevant unit is rarely the token.

It might be:

  • cost per booked appointment
  • cost per qualified lead
  • cost per completed interaction
  • cost per graded call
  • cost per resolved case
  • cost per successful automation

A dialer can make thousands of inexpensive attempts.

The client pays for the booked appointment.

The dial is not the business outcome.

Our equivalent unit for AI-assisted software development is a merged code change.

On Claude, the measured cost was approximately $1.00 per merged change under our subscription economics.

DeepSeek did not produce merged changes for us because the code-writing agents never went live.

So its corresponding figure has to be estimated.

Starting from the measured token economics, we estimated approximately $0.63 of DeepSeek API tokens per change off-peak, or around $0.76 when accounting for peak pricing.

We then applied assumptions for additional testing, review and retries, and for the possibility that the system would complete fewer jobs successfully.

Under those assumptions, the estimated cost per completed change ranged from approximately $1.15 to $4.90.

The middle estimate was approximately $2.40.

Those figures are not measurements.

They are scenario estimates based partly on the vendor-reported benchmark difference and assumptions about additional work.

That distinction matters.

The experiment measured the infrastructure and token economics.

It did not measure the cost of producing completed software changes with a production-ready DeepSeek coding system.

The Expertise Tax

There was another cost that does not appear neatly on a cloud invoice.

We encountered 14 separate configuration and operational traps, ranging from the nightly serving build to a progress indicator that reported completion before the model was actually ready.

A team running its own infrastructure must understand and manage these issues.

An API customer generally pays the provider to absorb that complexity.

This does not make self-hosting inherently wrong.

It means the comparison must include engineering labour, operational risk and security work alongside GPU rental.

For a large organisation with strong infrastructure capability and predictable, high utilisation, that equation may look different from the equation for a smaller team.

The workload itself matters just as much.

What Contact-Center Leaders Should Ask

The experiment was conducted on software engineering rather than customer-facing contact-center traffic.

That limitation should prevent anyone from treating these measurements as a direct forecast for a voice bot or QA deployment.

But the questions raised by the experiment apply directly to contact-center AI architecture.

A team evaluating self-hosted or privately deployed AI should ask at least three questions.

1. What proportion of every request is new versus cached?

A workload dominated by repeated context may have very different infrastructure economics from one in which every request contains substantial new information.

Quality assurance, for example, may repeatedly process new transcripts.

That could produce a very different cost profile from an engineering workload with heavy context reuse.

2. How busy will the infrastructure actually be?

A GPU that is fully utilised during business hours but largely idle overnight has a different economic profile from a workload that maintains high utilisation around the clock.

Capacity planning therefore needs to consider the complete operating day, not the peak benchmark.

3. Who owns the sandbox and surrounding security controls?

If an AI agent can access internal systems, customer information, tools or production environments, the model is only one component of the security boundary.

The orchestration layer, credentials, network controls, execution environment and monitoring become part of the deployment.

Those controls have an engineering cost.

What the Experiment Does—and Does Not—Prove

The experiment lasted roughly three hours.

The GPU provider reclaimed the spot machine at 14:08.

The team was back on Claude Opus 5.5 by approximately 14:15.

The total GPU cost was approximately $28.

The experiment does not prove that DeepSeek is more or less economical for contact-center AI generally.

It does not establish a controlled model-quality ranking.

It does not provide a benchmark of customer-facing voice agents.

Moreover, it does not establish the cost of production QA grading.

It does not eliminate the possibility that a different infrastructure configuration, utilisation pattern, workload or engineering team would produce different economics.

What it does provide is a detailed operator case study showing what happened when one contact center tried to run an open model on four rented H200 GPUs.

The results show the importance of looking beyond token prices.

They show how utilisation can change infrastructure economics.

They show that model serving can require significant operational expertise.

Moreover, they show that security controls around agents remain the operator’s responsibility.

And they show why business outcomes are a more useful unit of analysis than raw token consumption.

For a contact center, the ultimate question is not:

How cheap is a token?

It is:

What does it cost to produce the outcome the customer actually pays for?

Whether that outcome is a booked appointment, a completed interaction, a graded call or a resolved customer issue, the AI architecture should ultimately be evaluated against that unit.

That is where build-versus-buy decisions become meaningful.


AI Economics in the Contact Center: What a Four-H200 Test Revealed About Cost, Capacity and Risk

The Numbers at a Glance

MeasureValue
Infrastructure4 × NVIDIA H200, spot instance
Rental rate$9.186/hour
Test durationApproximately 3 hours 6 minutes
Approximate GPU cost$28
Successful configurationStart five
New-text throughput16,621 tokens/sec
Cached-text throughput521,027 tokens/sec
Written-output throughput5,281 tokens/sec
Estimated ceiling at our real workload mix~213 written tokens/sec
Estimated daily capacity at that mix~20 billion tokens/day
Median first responseClaude 3.5 sec; DeepSeek 12.8 sec
90th-percentile first responseClaude 9.6 sec; DeepSeek 56.6 sec
Terminal-Bench 4.0Vendor-reported: Claude 66.4; DeepSeek 31.2
Sep. 1–27 equivalent DeepSeek API cost~$3,517 off-peak; ~$7,034 peak
Claude subscription cost~$5,544
Same tokens at Opus 5.5 list API pricing~$154,600–$197,800
Estimated AI cost per merged changeClaude ~$1.00; DeepSeek estimate ~$1.15–$4.90

Method and Limitations

This was a one-day experiment on September 27, 2026, using one four-GPU H200 spot instance from one provider.

The experiment used internal software-engineering workloads and no customer data or customer calls.

Saturation tests ran for 60 seconds and were measured through vLLM metrics.

The daily capacity ceiling, estimated busy share and required GPU count are capacity-model estimates.

Latency figures came from transcript timestamps. The Claude and DeepSeek workloads differed in task, context size, tools and concurrency, so the latency figures represent operator experience rather than a controlled model benchmark.

Token volumes cover one development server from September 1 through September 27, 2026.

The token mix is based on the operator’s own sessions.

The per-change DeepSeek figures are estimates rather than measured production results.

The Terminal-Bench scores are vendor-reported and were not independently reproduced in this experiment.

The anonymised data sheet, including launch flags, saturation results, token volumes, cost calculations and per-call timing rows, is available from the authors on request.

Jason owns The Call Center Doctors, a US call center that provides outbound appointment setting, inbound answering, lead qualification and QA while building its own contact-center software in-house.

Original data-heavy case study: ccdocs.com/deepseek-lost/

Website: ccdocs.com

Related posts

Fractional AI teams: Revolutionizing enterprise innovation through India’s top 3% talent

Editor

Leeford Healthcare CX: Pharma with Quality and Care

Editor

Human AI Customer Experience Strategy: Designing CX for Performance, Empathy, and Scale

Editor

Leave a Comment