Business Case Executive Summary ES

AgentOps Hub

Telefónica Built by the Cloud Infrastructure & Operations team

AgentOps Hub is an AI-driven support automation platform for cloud engineering, built on Amazon Bedrock AgentCore and validated end to end. Specialized agents for domains like cost, observability, security, or remediation resolve engineering-grade L1 across four channels, and what escalates reaches L2 pre-digested: diagnosis, client context, and documentation attached.

5x
Volume Scaling
60%
Capacity Freed
24/7
Autonomous Ops
L1.5
Engineering-Grade Support

The innovation tax

Skilled engineers spend the majority of their time on repetitive, low-complexity tickets that follow predictable patterns: classifying, verifying, investigating, drafting. This "innovation tax" prevents scaling without proportional headcount growth.

Reactive Model
Teams respond to issues as they arrive, with no capacity for proactive service delivery or strategic customer engagement.
Linear Scaling
Growing the customer base requires proportionally growing headcount, a model that is neither sustainable nor economically viable.
Knowledge Silos
Complex issue resolutions stay locked with individuals rather than becoming reusable, searchable team assets.
Response Limits
Human-dependent support cannot deliver 24/7 near-instantaneous responses at scale without massive operational cost.
Talent Misallocation
Engineers capable of architectural guidance spend their time on routine L1 ticket resolution rather than high-value work.
Scalability Ceiling
Current infrastructure cannot handle 5x ticket volume growth without major investment in headcount and tooling.

From MVPs to Platform

Two independent MVPs, Cost Assistant and Observability Agent, built entirely by the Cloud Infrastructure & Operations team and validated with 3 select clients in Q1 2026 with 80% faster information retrieval. That foundation is now a unified multi-agent platform, built and validated end to end in the cloud.

80%
Faster Info Retrieval (Validated)
2 → 1
MVPs Unified into One Platform
Key Design Principles
✓ Three operating modes per client and per action: shadow (the AI works in the background while operators calibrate its work), assisted (the AI proposes a reply and the operator decides whether to edit it or approve it as the ticket response), and autonomous (the AI works autonomously, replying to tickets on its own)
✓ Human-in-the-loop: any action executed on the client platform must be approved by an operator
✓ Audited by design: from the platform's own operation to the step-by-step of every ticket or call, with a full trace of every decision, tool, and cost
✓ Domain-specific specialized agents, not one-size-fits-all
✓ Every resolution feeds a per-client knowledge base that grounds the next diagnosis
✓ AgentCore Memory (STM + LTM): agents adapt to each client with near-human contextual understanding, no manual parametrization
✓ Managed without code: prompts, agents, and runbooks are edited in the console and propagate live to the runtime
Validated Scenarios
L1 ticket, end to end An incident arrives from the ITSM and the supervisor triages it, asking the requester for more information right on the ticket if something is missing. With full context it routes the case to the right specialist agent, which diagnoses it with real account data and posts the reply. Everything is recorded: full trace and cost per ticket in the console.
Verified remediation A runbook, custom or an official AWS SSM Automation document, is proposed with real pre-checks, executed only after explicit approval, and verified afterwards: the fix is confirmed, not assumed.
Proactive detection A CloudWatch alarm, a Security Hub finding, or an AWS Health event is detected, diagnosed, and lands in the approval inbox as a draft ticket. If the severity is critical, the agent goes one step further: it places a call through Amazon Connect to the on-call engineer, telling them what is happening and that a ticket is already drafted and waiting for approval.
Asking the customer's AWS DevOps Agent If the specialist finds it lacks the information to propose a solid solution, it contacts the customer's AWS DevOps Agent, agent to agent (A2A), and commissions an in-depth study of their platform. The result comes back as input for the diagnosis: knowledge only the customer's own agent has, folded into the answer.

Specialized AI agents

A supervisor agent triages each request, gathers context, and routes to the right specialist agent. Specialists reach AWS through a governed tool catalog on AgentCore Gateway, with least-privilege cross-account access, and every write action goes through the human approval inbox. And when a ticket does need a human, it never arrives raw: L2 receives the investigation already done.

Supervisor
Supervisor Agent
Analyzes intent, validates completeness, requests clarification if needed, and routes to the right specialist agent.
Observability
Observability Agent
CloudWatch metrics, logs, and alarms. X-Ray distributed traces. Guided troubleshooting flow: alarms → metrics → logs → traces.
Cost
Cost Agent
AWS cost queries, forecasting, anomaly detection, and rightsizing recommendations. Access to Cost Explorer, Compute Optimizer, Budgets, and Pricing APIs.
Security
Security & Compliance Agent
Security posture and findings across your AWS accounts: AWS Security Hub when available, with graceful fallback to native checks. Critical findings feed the proactive flow.
Remediation
Remediation Agent
Executes curated runbooks, both custom and official AWS SSM Automation documents, with real pre-checks, deterministic guards, explicit human approval, and post-execution verification.
DevOps Agent
AWS DevOps Agent · A2A
Agent-to-agent integration with AWS's own DevOps Agent: when a specialist needs knowledge only the customer's agent has (infrastructure, deployments, history), it asks it directly.

Multi-channel serverless architecture

Access the platform from any of its four channels: web chat, phone call, ITSM ticket, or the proactive channel watching your accounts. The same supervisor handles every request, with AgentCore Runtime, Memory, and Gateway as the serverless foundation, all governed live from a single control-plane console.

Channel 1 · Web Chat
User sends message → Supervisor analyzes intent → Routes to specialist → AI drafts response → User reviews
Channel 2 · Voice (Amazon Connect)
User calls → Voice AI transcribes → Supervisor routes → Specialist resolves → Voice AI responds
Channel 3 · ITSM (ITSMTE bus)
Ticket created → Supervisor triages → Specialist investigates → Operator approves → Reply posted to the ticket
Channel 4 · Proactive (alarms · findings · AWS Health)
Signal detected → Specialist diagnoses → Draft ticket to inbox → Operator approves → On-call paged if critical
Amazon Bedrock
Best-fit model per agent · Guardrails
AgentCore
Runtime · Memory · Gateway
DynamoDB + KMS
Encrypted · On-demand
CloudWatch
Logs, Metrics, Alarms

From foundation to production

From hands-on experimentation to client validation to enterprise production: a journey driven by curiosity, confirmed by real results.

December 2025
First MVPs
Cost Assistant and Observability Agent built independently after AWS training events. Validated with AgentCore Runtime, Strands SDK, and Bedrock from day one.
Q1 2026
Validation & Strategic Vision
Presented to leadership. Deployed to 3 select clients, with 80% faster information retrieval. AWS training on multi-agent orchestration confirms the path. Decision to unify under a single supervisor with ITSM integration.
April – May 2026
Architecture & Foundation
Rebuilding from the ground up: unifying both MVPs under a single supervisor, designing voice flows, and laying the groundwork for plug-and-play agent onboarding.
May – August 2026
Delivery & Build
Platform built and validated end to end in the cloud: four specialists with Lambda tools on AgentCore Gateway, ITSM channel on the real ITSMTE contract, web chat with memory, proactive flow, approval inbox with shadow, assisted, and autonomous modes, verified runbooks, AWS escalation (Support cases and quota requests), and per-ticket observability with cost.
September 2026
Pilot & Hardening
Onboarding the first real clients in shadow mode, measuring deflection, latency, and quality per client before graduating actions to assisted and autonomous.
October 2026
Go-Live
Final validation across all channels, security review, and production go-live for our managed clients as a fully operational service.
Future Opportunities
AI FinOps The platform already measures the cost of every ticket and reports savings weekly (caching, right-sized models). Next: applying those optimizations automatically, so the AI that runs support also optimizes its own spend.
Self-improving agents An LLM judge already reviews each specialist's quality every week. Next: agents that tune their own prompts from operator feedback and real outcomes, with every change versioned and approved by a human.
Total Duration
10
Months, MVPs to Production
Sep 2026
Pilot Starts
Oct 2026
Go-Live

Investment model

Strategic co-investment: Telefónica contributes specialized engineering, client knowledge, and product vision; AWS provides the managed AI platform, cloud infrastructure, and specialized training. Aligned incentives, shared risk, serverless economics that scale with demand.

Co-investment
Strategic Partnership
Talent + Platform
Telefónica + AWS
GenAI Operations
Next-gen AWS AI Services
Cost distribution by service
Infrastructure specs
RegionEurope (Frankfurt) · eu-central-1
Bedrock InferenceCross-Region, On Demand
Tokens≈50,000 per resolved ticket
S3 StorageKnowledge bases + static assets
DynamoDBOn-demand · KMS-encrypted
Lambda + FargateServerless compute
ObservabilityFull trace per ticket · OTel
Data TransferAPI-level integrations, minimal egress

Operational transformation

From reactive to proactive. From knowledge silos to compounding team assets. From linear scaling to exponential capacity.

Metric Without AgentOps AgentOps Assisted AgentOps Autonomous
L1 Response Time Minutes to hours Seconds (AI analysis) Seconds, 24/7
Support Availability 12x5 + on-call AI 24/7 · approvals 12x5 + on-call 24/7/365 autonomous
Volume Capacity Baseline +40–60% 5x baseline
Engineer Time on L1 ~100% Reduced (draft review) ~40% (60% freed)
Escalation Quality Raw ticket forwarded to L2 Pre-digested: diagnosis, context, docs Pre-digested + vendor escalation ready
Knowledge Capture Individual / ad-hoc Per-client KB, fed by every resolution + full client ITSM history, curated
Operating Model Reactive Semi-automated Fully proactive
For Support Teams
Analysis in seconds. Escalations arrive pre-digested. Knowledge builds automatically. Focus on the work that truly needs an engineer.
For Customers
Faster resolution. Expert-level analysis. Self-service knowledge library. Consistent quality across all interactions.
For the Business
Scale 5x without proportional cost. AI-powered differentiation. Knowledge compounds over time. Serverless economics.

Automate the routine.
Capture the knowledge.
Unleash the talent.

The foundation is proven. The platform is built and validated end to end. The pilot is next. The business case is clear.

Dec 2025
Foundation Validated ✓
Q1 2026
Enhancements Completed ✓
Aug 2026
Platform Built ✓
Oct 2026
Go-Live

The people making it happen

A joint initiative driven by Cloud Infrastructure & Operations engineering, sponsored by Telefónica leadership, and supported by AWS Partner Technical Account Management.

Cloud Infrastructure & Operations · Drivers
Tech Lead
Maykel Cano Serrano
Maykel Cano Serrano
Head of Cloud Infrastructure
Cloud Infrastructure & Operations, Telefónica
LinkedIn
Marcos Moreno Torres
Marcos Moreno Torres
Senior Cloud Architect
Cloud Infrastructure & Operations, Telefónica
LinkedIn
Telefónica · Sponsors
Jordi Costa Mur
Jordi Costa Mur
Head of Hybrid Multicloud & Hyperscalers
Telefónica
LinkedIn
Álvaro Paniagua Martín
Álvaro Paniagua Martín
Chief Operating Officer (COO)
Telefónica
LinkedIn
Amazon Web Services
Oriol Cañellas Tortell
Oriol Cañellas Tortell
AWS Partner Technical Account Manager
Amazon Web Services
LinkedIn

Dive deeper

Pick the format that matches your moment. A quick executive snapshot to share with leadership, or the full business case with architecture, costs, risks, and the complete project plan.