01 · Executive Summary
A next-generation support automation initiative
AgentOps Hub is a next-generation, AI-powered support automation initiative. Born from hands-on experimentation with AWS AgentCore after training events in December 2025, it evolved from two independent MVPs (a Cost Assistant and an Observability Agent) into a unified, production-grade Tier 1 support automation platform.
Both MVPs were validated with 3 select clients in Q1 2026, achieving an 80% reduction in information retrieval time. AgentOps Hub takes this proven foundation to enterprise scale with a multi-channel serverless architecture (web chat, voice, ITSM, and proactive detection) built on Amazon Bedrock AgentCore, targeting 5x volume capacity, 60% capacity reallocation toward innovation, and 24/7 autonomous operation. As of August 2026 the platform is built and validated end to end in the cloud. The result is what we call Level 1.5 support: engineering-grade L1 resolved end to end, with escalations that reach L2 pre-digested.
Total infrastructure investment covered through a strategic co-investment with AWS (Telefónica: talent · AWS: platform).
Timeline: First MVPs (December 2025), client validation & strategic vision (Q1 2026), full production deployment April – October 2026.
02 · Problem Statement
The current state
Partner support teams manage expanding portfolios of technical inquiries across increasingly complex AWS environments. Each ticket follows a familiar journey: classify, verify information, investigate resources, draft guidance, monitor responses, and coordinate escalations. Technical engineers are heavily burdened by these repetitive, low-complexity tasks that do not require human intuition, creating a significant "innovation tax."
Key challenges
- Reactive support model: Teams respond to issues as they arrive, with no capacity for proactive service.
- Single-channel constraint: Customers are funneled through one interface (ITSM tickets) even when a quick chat or phone call would resolve the issue faster.
- Linear scaling constraint: Growing the customer base requires proportionally growing headcount.
- Talent misallocation: Engineers capable of complex troubleshooting spend time on routine analysis.
- Knowledge silos: Complex issue resolutions stay locked with individuals rather than becoming team assets.
- Response time limitations: Human-dependent support cannot deliver 24/7 near-instantaneous responses.
- Scalability ceiling: Current infrastructure cannot handle 5x ticket volume growth.
Business impact of inaction
- Inability to scale the customer base without significant hiring investment.
- Continued erosion of engineering capacity for innovation and proactive client improvements.
- Loss of institutional knowledge as it remains locked with individual engineers.
- Competitive disadvantage as peers adopt AI-driven support models.
- Increasing operational costs per ticket as complexity and volume grow.
03 · Strategic Vision
From proof-of-concept to production scale
A deliberate two-phase strategy validates the concept, then scales to enterprise production.
Phase A: First MVPs & Validation (December 2025 – Q1 2026)
Two independent MVPs built after AWS training events on AgentCore and Strands Agents SDK:
- Cost Assistant (Maykel Cano) and Observability Agent (Marcos Moreno), standalone web apps with chat interface connected to client AWS accounts.
- Built on AgentCore Runtime, Strands SDK, Bedrock (Claude), Cognito with MFA, DynamoDB from day one.
- Presented to Telefónica leadership. Deployed to 3 select clients, 80% faster information retrieval for decision-making.
- AWS training on multi-agent orchestration confirms the unified platform approach: the demo featured 3 agents (cost, observability, remediation), 2 of which the team had independently built.
Phase B: AgentOps Hub (April – October 2026)
Takes the validated foundation to enterprise production scale with a multi-channel serverless architecture:
- Four entry channels (Web Chat on CloudFront, Voice with Amazon Connect + Nova Sonic S2S, ITSM via Telefónica's ITSMTE bus, and Proactive detection), all converging on the same Supervisor.
- Bedrock AgentCore for managed agent runtime, memory, and gateway.
- Lambda tools on AgentCore Gateway: a versioned, governed tool catalog shared by all specialist agents.
- Human approval inbox with shadow, assisted, and autonomous modes per client and per action.
- Client knowledge: a per-client profile plus a Bedrock Knowledge Base fed by every resolution.
- Control-plane console: prompts, agents, runbooks, users (RBAC), alerts, per-ticket traces, KPIs, and cost per incident.
- Multi-account support, AWS Support and Service Quotas escalation, AWS DevOps Agent integration (A2A), cross-region inference.
Strategic objectives
| # | Objective | Target |
| 1 | Automate Response | Near-instantaneous response times for L1 inquiries using specialized AI agents and Amazon Bedrock. |
| 2 | Multi-Channel Access | Web chat, voice (phone), and ITSM tickets, all served by the same supervisor and knowledge base. |
| 3 | Enhance Scalability | Serverless backend capable of handling 5x current ticket volume. |
| 4 | Proactive Value | Reallocate 60% of support team capacity toward proactive consulting and new solution development. |
| 5 | Capture Knowledge | Automatic documentation builds a searchable, ever-growing knowledge base from every resolved interaction. |
| 6 | Full Observability | Real-time monitoring of AI accuracy and system performance via Amazon CloudWatch. |
04 · Validated Foundation
Multi-agent architecture
| Agent | Domain Expertise |
| Supervisor Agent | Validates AWS context, extracts service & region info, routes to the most appropriate specialist agent. |
| Observability Agent | CloudWatch metrics, logs, and alarms. X-Ray distributed traces. Guided troubleshooting: alarms → metrics → logs → traces. |
| Cost Agent | AWS cost queries, forecasting, anomaly detection, rightsizing. Access to Cost Explorer, Compute Optimizer, Budgets, Pricing. |
| Security & Compliance Agent | Security posture and findings: AWS Security Hub when the client has it, native checks otherwise. Critical findings feed the proactive flow. |
| Remediation Agent | Curated runbooks (custom and official AWS SSM Automation documents) with real pre-checks, deterministic guards, human approval, and verified execution. |
| AWS DevOps Agent (A2A) | Agent-to-agent integration: customer-specific knowledge (infrastructure, deployments, history) the specialists consult when their own tools are not enough. |
Validated real-world scenarios
- Cost Analysis: "How much did we spend this quarter on EC2?": the Supervisor routes to the Cost Agent, which queries Cost Explorer, breaks down by service, account, and region, and provides optimization recommendations.
- Active Alarms: "What are the active alarms in production?": the Supervisor routes to the Observability Agent, which checks CloudWatch alarms and provides guided troubleshooting: alarms → metrics → logs → traces.
- Log Analysis: "Show me error logs from the last hour": the Supervisor routes to the Observability Agent for CloudWatch log analysis with pattern detection and root cause identification.
05 · Solution Architecture
Multi-channel serverless architecture
Four entry channels (web chat, voice, ITSM, and proactive detection) converge on the same Supervisor, the same specialist agents, and the same knowledge. Every component is serverless; every interaction is isolated; every channel benefits from every other channel's learnings.
Entry channels
| Channel | Flow |
| 1 · Web Chat | User sends message → Supervisor analyzes intent → Routes to specialist → AI drafts response → User reviews. |
| 2 · Voice | User calls → Voice AI transcribes → Supervisor routes → Specialist resolves → Voice AI responds. |
| 3 · ITSM | Ticket created → Supervisor triages → Specialist investigates → Operator approves → Reply posted to the ticket. |
| 4 · Proactive | Alarm, finding, or AWS Health event detected → Specialist diagnoses → Draft ticket to the inbox → Operator approves → On-call paged if critical. |
Platform layers
| Layer | Components |
| Entry Points | CloudFront (Web Chat), Amazon Connect (Voice), ITSMTE bus poller (ITSM), proactive detectors (alarms, findings, AWS Health). |
| AI / Agent Layer | AgentCore Runtime, Gateway, Memory; the leading models on Amazon Bedrock (Claude, Nova, GPT-OSS); Nova Sonic speech-to-speech. |
| Specialized Agents | Supervisor, Observability, Cost, Security & Compliance, Remediation + AWS DevOps Agent via A2A. |
| Compute | ECS Fargate, AWS Lambda. |
| Knowledge & Storage | Amazon S3, Amazon DynamoDB, Bedrock Knowledge Bases. |
| Observability | Per-ticket traces with cost, KPIs, Amazon CloudWatch, AWS X-Ray (OpenTelemetry). |
| Integration | ITSM (ITSMTE bus), AWS Support & Service Quotas, AWS DevOps Agent, AWS SSM Automation, Bedrock Knowledge Bases. |
AWS services inventory
| Service | Configuration | Purpose |
| Amazon Bedrock | Cross-Region inference, On Demand. | Generative AI inference for all agents. |
| Bedrock AgentCore | Runtime, Gateway, Memory. | Managed agent orchestration & observability. |
| Amazon Connect | AI agents · Nova Sonic S2S · Return-to-Control. | Voice channel: bidirectional audio streaming with human escalation. |
| Amazon CloudFront | Global CDN in front of the React SPA. | Low-latency console and web chat delivery. |
| ALB + ECS Fargate | Containerized FastAPI backend. | Control-plane API, channel pollers, evaluators. |
| AWS Lambda | Versioned tool targets. | Tool catalog behind AgentCore Gateway (cost, observability, security, remediation, escalation, knowledge). |
| DynamoDB | On-demand, KMS encryption. | Conversations, session state, account configuration. |
| S3 | Standard storage. | Frontend hosting (SPA), data source for Bedrock Knowledge Bases. |
| CloudWatch | Logs, Metrics, Alarms, Dashboards. | Full observability stack. |
| Secrets Manager | Encrypted credential storage. | API keys, integration tokens, service credentials. |
Architecture design patterns
- Multi-Channel Access: Web chat, voice, ITSM tickets, and proactive detections all converge on the same Supervisor via AgentCore Runtime and Gateway, so every channel reuses the same agents and knowledge.
- Multi-Agent Specialization: Supervisor routes to domain-specific agents rather than a single general-purpose AI.
- Lambda Tools: Specialists reach AWS through versioned Lambda targets on AgentCore Gateway: one governed, reusable tool catalog.
- Human-in-the-Loop: Every write goes through the approval inbox; actions graduate to autonomous on real data and degrade automatically on negative feedback.
- Serverless-First: Every compute component scales automatically.
- Knowledge-Augmented AI: per-client profile plus Bedrock Knowledge Bases, continuously enriched from resolved interactions.
- Full Observability: Per-ticket traces with tool calls, durations, and cost, plus KPIs, CloudWatch, and X-Ray.
- Agent Isolation: AgentCore provides runtime isolation, with strict per-organization data isolation across the platform.
- ITSM Integration: bidirectional integration with the ITSMTE bus, safe against duplicates and loops; every reply goes through the approval flow and is fully audited.
06 · Cost Analysis
Investment model
A strategic co-investment where each partner contributes their core strength. Telefónica brings the engineering talent, client relationships, and product vision. AWS provides the managed AI platform, cloud infrastructure, and specialized training through the AI Strategic Initiative.
Co-investment
Strategic Partnership
Aligned incentives: Telefónica scales its managed services offering with AI-powered differentiation; AWS grows platform adoption through a production reference architecture. Shared risk, shared upside, serverless economics that scale with demand.
Cost distribution by service
| Service | % of Total |
| Amazon Bedrock (AI models) | 69% |
| Amazon Bedrock AgentCore | 9% |
| ECS Fargate + ALB | 8.5% |
| Amazon CloudWatch · X-Ray | 7% |
| Amazon DynamoDB | 4% |
| Amazon S3 + CloudFront | 2% |
| AWS Lambda | 0.5% |
07 · Project Plan
Full timeline
| Phase | Duration | Start | End | Status |
| First MVPs | n/a | 2025 | Dec 2025 | Validated |
| Q1 2026 Enhancements | ~90 days | Jan 2026 | Mar 2026 | Completed |
| Architecture Design | 30 days | 20 Apr 2026 | 20 May 2026 | Completed |
| Delivery / Build | 100 days | 20 May 2026 | Aug 2026 | Completed |
| Pilot & Hardening | ~30 days | Sep 2026 | 30 Sep 2026 | In Progress |
| Testing & Go-Live | ~15 days | 30 Sep 2026 | 16 Oct 2026 | Planned |
Phase 1: Architecture Design (30 days)
Rebuild from the ground up: unify both MVPs under a single supervisor, design Amazon Connect voice flows, define API contracts for all agents, establish DynamoDB data models, configure AgentCore Memory and Gateway, and lay the groundwork for plug-and-play agent onboarding.
Phase 2: Delivery / Build (100 days)
End-to-end orchestration with multi-account AWS support, delivered and validated in the cloud: ITSM channel on the real ITSMTE contract, web chat with memory, proactive flow, approval inbox with shadow, assisted, and autonomous modes, verified runbooks (custom and AWS SSM Automation), AWS Support and Quotas escalation, client knowledge base, and per-ticket observability with cost.
Phase 3: Testing & Validation (50 days)
End-to-end validation across all channels with real scenarios, AI accuracy testing against MVP baseline, load testing, multi-account security review, observability verification, and production deployment for managed clients.
08 · Business Value
Operational transformation
| Metric | Without AgentOps | AgentOps Assisted | AgentOps Autonomous |
| L1 Response Time | Minutes to hours | Seconds (AI analysis) | Seconds, 24/7 |
| Support Availability | 12x5 + on-call | AI 24/7 · approvals 12x5 + on-call | 24/7/365 autonomous |
| Volume Capacity | Baseline | +40–60% | 5x baseline |
| Engineer Time on L1 | ~100% | Reduced (draft review) | ~40% (60% freed) |
| Escalation Quality | Raw ticket forwarded to L2 | Pre-digested: diagnosis, context, docs | Pre-digested + vendor escalation ready |
| Knowledge Capture | Individual / ad-hoc | Per-client KB, fed by every resolution | + full client ITSM history, curated |
| Operating Model | Reactive | Semi-automated | Fully proactive |
Value for the business
- Multi-Channel Customer Experience: Customers choose how they engage (web chat, phone call, or ticket) with consistent quality and shared context across all channels.
- Level 1.5 Handovers: tickets that do escalate reach L2 with the diagnosis, client context, and documentation already attached, cutting resolution time on the work AI does not close.
- Revenue Growth: Scale customer base 5x without proportional cost increase.
- Talent Optimization: 60% of support engineering capacity redirected to proactive consulting, innovation, and feature development.
- Competitive Differentiation: AI-powered support with specialized agents and voice-first interaction differentiates the offering.
- Knowledge Compounding: Every resolved interaction, regardless of channel, enriches the knowledge base, making future resolutions faster.
- Client Retention: Faster response times and proactive service improve satisfaction and reduce churn.
- Serverless Economics: Pay-per-use infrastructure scales automatically with demand.
09 · Risk Considerations
Risks & mitigations
| Risk | Mitigation |
| AI response accuracy | Human approval inbox for every write, per-ticket traces with quality scores, and autonomous actions that degrade back to assisted on negative feedback. |
| Cost overrun on Bedrock | Cost per incident measured on every ticket, editable pricing, weekly optimization reports (caching, model right-sizing), cross-region inference. |
| Knowledge base staleness | Bedrock Knowledge Bases auto-enriched from every resolution, fully deletable per client, version-controlled prompts. |
| Multi-account security | Least-privilege cross-account roles, with write access as separate opt-in stacks per capability; Secrets Manager; strict per-organization isolation. |
| Integration complexity | Real ITSMTE contract received and implemented, validated against a live mock, safe against duplicates and loops, with complete audit logging. |
| New channels (voice, web) | Amazon Connect + Nova Sonic are managed services; Return-to-Control provides a proven escalation path to human agents; channels share the same supervisor, minimizing divergence. |
| Timeline risk (100-day build) | Mitigated: the build completed in August 2026, validated end to end in the cloud. |
| Agent routing accuracy | LLM triage with a conservative rule (escalate on doubt), routing measured per ticket in the console, continuous improvement through live prompt versioning. |
10 · Conclusion
A strategic investment
AgentOps Hub represents a strategic investment in AI-driven operational transformation, built on a proven foundation. Two independent MVPs (Cost Assistant and Observability Agent) have already validated that specialized AI agents can deliver 80% faster information retrieval for real clients, with AgentCore, Strands SDK, and Bedrock from day one. The production architecture unifies these validated agents into a full multi-channel serverless platform (web chat, voice, ITSM, and proactive detection), all sharing the same supervisor and knowledge base.
Through a strategic co-investment with AWS (Telefónica contributing talent, AWS providing the platform), the production deployment delivers 5x support capacity across every channel, 60% talent reallocation, automatic knowledge capture, and a fundamental shift from reactive to proactive service delivery.
The foundation is proven. The platform is built and validated end to end. The pilot is next. The business case is clear: automate the routine, meet customers on every channel, capture the knowledge, unleash the talent, and scale without limits.