How NVIDIA is Using Agentic AI to Build Autonomous Telecom Networks

Telcos are moving beyond predefined automation toward agentic AI that can reason, research, optimize, and safely operate networks.

How NVIDIA is Using Agentic AI to Build Autonomous Telecom Networks
Rohit Lakhotia

Share

Telecom networks are already heavily automated. When a known problem occurs, systems can follow predefined procedures, execute configurations, and keep many day-to-day operations moving without constant human intervention. But automation has a limit: it can only execute what someone has already defined.

The harder question is what happens when the network encounters a problem that has no existing runbook or when the existing solution works, but there might be a better one for reducing cost, improving latency, increasing resilience, or saving energy. This is where the idea of agentic AI starts to become interesting.

Instead of simply executing a predefined procedure, an autonomous agent can understand an operator's intent, observe the network, investigate what is happening, develop and compare possible plans, and coordinate actions across different systems while still operating within defined policies and governance.

That's the shift NVIDIA is exploring: moving telecom networks from automation that executes known solutions toward autonomy that can reason about problems, discover solutions, and safely act on them.

And getting there isn't just about choosing a more capable AI model. Telcos need an underlying autonomy platform that brings together telecom-domain models, data, tools, digital twins, agent harnesses, and secure execution environments.

Let’s discover how NVIDIA did this in this blog today.

The Problem–Solution Loop

Think about a network incident. An operator might receive a customer ticket. Or the network might detect an anomaly through telemetry.

The first question is: Have we seen this problem before?

If the answer is yes, the system can follow an execute path. An existing reasoning trace, runbook, or trusted skill can be used to determine what needs to happen. An on-demand agent can execute the task, while a long-running agent can continuously monitor the situation if the solution needs to be applied and verified over time.

But there is another possibility. Maybe the problem is understood, but the existing solution isn't necessarily the best one. That becomes the optimize path. For example, operators may want to improve the network against measurable objectives such as energy efficiency, latency, resilience, or cost. Deep-research agents can explore alternatives and produce ranked optimization plans. A long-running agent can then apply the selected plan, monitor the results, and re-optimize or roll back when necessary.

And then there is the hardest case. The problem has never been encountered before. There is no existing reasoning trace that matches it. This becomes the discovery path. Deep-research agents can investigate the unfamiliar behavior, correlate signals across different domains, and turn the unknown situation into a well-defined problem. From there, other agents can take discrete actions or manage longer-running recovery.

Over time, something interesting happens. The solutions and execution traces discovered during these investigations can be turned into new or updated skills. So a problem that originally required research can eventually become a governed execution path. The autonomy library effectively grows as the system learns new ways of solving problems.

Not All Agents Do the Same Job

NVIDIA describes three important types of agents in this model.

  1. On-demand agents handle bounded tasks. They might apply configuration changes, run NOC scripts, or answer customer-care questions.

  2. Long-running agents stay with a problem for a longer period. They continuously sense the network, validate actions, coordinate across systems, and decide when to escalate, roll back, or re-optimize.

  3. Deep-research agents, these are used when the answer isn't already known. Instead of returning one fixed solution, they explore different data sources, tools, and digital twins to propose, validate, and rank alternative plans.

This distinction is important because autonomous networks aren't simply about deploying one giant AI agent. Different agents have different responsibilities, and they need to work together around the same problem–solution loop.

That brings us to the bigger question: What does the platform underneath all these agents actually look like?

The Platform Behind Autonomous Telecom Agents

If every telecom use case builds its own agent, its own tools, its own policies, and its own execution environment, telcos would simply end up with another collection of disconnected systems. The goal instead is to create a shared autonomy platform.

At the center are telecom agents that understand network and service behavior and can turn that understanding into closed-loop actions. Around them sit four important building blocks:

Data and models → Agent harnesses → Tools and digital twins → Secure runtime

Together, these provide the foundation that agents need to reason, act, and operate safely.

Data and models

The first layer is the data. Telecom agents need high-quality network and customer data to understand what is happening in the environment.

NVIDIA describes using NeMo Data Designer and NeMo Safe Synthesizer to generate synthetic data and anonymize sensitive records. This can increase the volume and diversity of production-like datasets while preserving privacy.

Reasoning models such as Nemotron can then be fine-tuned on these datasets and grounded in telecom ontologies and operational context. The goal isn't simply to make a model generate text. It needs to interpret network signals, form hypotheses, validate them, and reason about why a particular sequence of actions, tool calls, and decisions is safe and effective.

NVIDIA also describes NV-Tesseract time-series models for analyzing multivariate network telemetry, detecting anomalies, and forecasting network behavior. These signals can then support proactive detection and remediation workflows.

The Agent Harness

A model by itself isn't an autonomous agent. The agent harness provides the control loop around the model. It receives the intent, manages session state and memory, determines when additional context is required, chooses which tools or digital twins to use, and decides when specialized skills such as deep research should be invoked.

NVIDIA's Agent Toolkit provides building blocks for connecting these agent harnesses to shared tools, observability, and evaluation frameworks. This makes it possible to deploy and orchestrate telecom agent workflows more reliably rather than building every workflow from scratch.

But telecom networks have another requirement that is arguably even more important: Agents cannot simply be given unrestricted access to production systems.

The Secure Runtime

Telecom networks operate under strict reliability and regulatory constraints. So autonomous agents need clearly enforced security and governance boundaries.

NVIDIA's OpenShell provides isolated sandboxes for agents and controls their access to filesystems, networks, tools, and inference endpoints according to corporate policies.

NemoClaw provides the blueprint for agent deployment, lifecycle management, and policy rollout. This gives the autonomy platform a controlled environment in which agents can actually operate. And when these layers come together, the goal is bigger than any individual use case.

Instead of creating one bespoke agent for anomaly detection, another for customer care, and another for network optimization, they can all use the same telecom-aware reasoning foundations, tools, and secure runtime.

What Does Deep Research Actually Look Like?

This is where the difference between traditional automation and agentic AI becomes much clearer. Traditional automation starts with a known procedure. Deep research starts with a problem where the answer may not be obvious.

NVIDIA's AI-Q provides an example of this approach. A planner agent first frames the problem and determines which domains and data sources are relevant. Then researcher agents can fan out across OSS/BSS systems, telemetry, and digital twins to collect evidence in parallel.

An orchestrator brings those findings together and can drive additional research passes until the required quality and risk thresholds are reached.

The result isn't simply: "Do X."

Instead, the system produces a ranked set of proposals, tied back to the underlying data and simulations. Those proposals can then be passed to agents that apply changes under policy, monitor what happens afterward, and trigger fallbacks or additional research if the expected targets aren't achieved. And for higher-risk domains, NVIDIA notes that these loops can have explicit approval thresholds so that operators review proposals before production changes are made.

From the Platform to Real Network Problems

One example NVIDIA discusses is autonomous anomaly detection and remediation in carrier-grade SR-MPLS networks.

When telemetry indicates congestion, tunnel degradation, or link failures, a deep-research agent can examine topology and routing state, analyze performance metrics, and compare alternative SR-TE paths or routing policies. Instead of returning a single remediation, it produces ranked options along with trade-offs involving performance, risk, and policy.

A long-running agent then acts as the execution spine. It chooses a plan, coordinates the required steps across SDN controllers and traffic-engineering tools, and watches the resulting telemetry. If the network doesn't recover as expected, it can fall back to another plan.

The environment can also serve as a testbed, allowing teams to generate structured traces, fine-tune telecom reasoning models, and validate new autonomy patterns before bringing them closer to production.

But NVIDIA's examples don't stop at network operations. Agentic AI can also be used to design new network algorithms.

When Agents Start Designing Network Algorithms

NVIDIA Research's AI Telco Engineer takes a wireless PHY- or MAC-layer problem together with a scoring function and uses agentic evolutionary search to discover algorithms that can meet or outperform established baselines.

The process works somewhat like an evolutionary search. A meta agent proposes different algorithm ideas. Parallel agents implement and evaluate those ideas using tools such as Sionna, NVIDIA's GPU-accelerated wireless simulation library. The better-performing ideas are retained, combined, and developed further in later generations, while new ideas continue to be explored.

In early experiments, NVIDIA reports that the AI Telco Engineer generated explainable PHY/MAC-layer algorithms that matched strong classical methods on channel-estimation and achieved more than a 3% spectral-efficiency gain over the industry-standard solution for link adaptation. This illustrates the broader direction NVIDIA is describing. The role of agents doesn't have to stop at operating the network. They can also help discover and develop new ways of operating it.

So, What Does an AI-Native Telco Look Like?

The important shift isn't simply putting an AI model inside a telecom workflow. It's building a platform where agents can repeatedly move through the complete loop:

Intent or event → Understand the problem → Research or retrieve knowledge → Develop plans → Validate → Execute under policy → Observe the result → Re-optimize, roll back, or learn

Some problems will follow an existing execution path. Others will require optimization. And completely unfamiliar problems may require discovery. As those discoveries become reusable skills, the system can gradually expand the set of problems that can be handled through governed execution. That is the bigger idea behind NVIDIA's autonomy platform.

Instead of treating agents as isolated experiments, telcos can make them the first users of a shared platform containing telecom reasoning models, ontologies, tools, simulations, policies, and secure execution environments. Each new use case can then strengthen the same underlying stack rather than creating another siloed automation system.

And that is what moves the conversation from AI-assisted network operations toward autonomous networks.

Key Takeaways

  • Today's telecom automation is largely focused on executing predefined solutions, typically around Level 2–3 autonomy.

  • Higher autonomy requires agents that can understand intent, reason about network state, research solutions, plan, execute, and monitor outcomes.

  • NVIDIA describes three agent types: on-demand, long-running, and deep-research agents.

  • Problems can follow three paths: execute, optimize, or discover.

  • New research outcomes and execution traces can become reusable skills.

  • A telco autonomy platform needs shared data, telecom-domain models, agent harnesses, tools, digital twins, and secure runtime infrastructure.

  • Deep-research agents can produce ranked and validated plans instead of one-shot fixes.

  • Long-running agents can close the loop by executing plans, monitoring results, and rolling back or re-optimizing when needed.

  • The same agentic approach can extend beyond network operations into network algorithm discovery and design.

Official blog from Nvidia: How Telcos Build Autonomous Networks with Agentic AI

By now, you must have had a clear idea of, How NVIDIA is Using Agentic AI to Build Autonomous Telecom Networks? In a nutshell, NVIDIA explores how telcos can move from predefined automation to agentic AI that can reason, research, plan, execute, and monitor network operations. The goal is to build a shared autonomy platform where agents work with telecom models, tools, digital twins, and secure runtimes to solve both known and unfamiliar network problems.

Congratulations! You've just advanced another step in your tech journey. Keep progressing!

Share