Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- The evolution of IT automation: from static runbooks to reasoning agents
- Agent anatomy: exploring the reasoning loop, tool usage, memory, and planning
- Determining when to automate versus when to retain human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm patterns
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Building your first operational agent: querying monitoring, diagnosing, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and routing
- Generating root cause hypotheses and gathering evidence
- Automated remediation: restarts, scaling, rollbacks, and failover actions
- Creating an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Action classification: read-only, low-risk, high-risk, and destructive
- Establishing approval gates and escalation policies for critical operations
- Guardrail patterns: action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent response
Observability and Evaluation
- Tracing agent reasoning chains for debugging and audit purposes
- Assessing agent decision quality: precision, recall, and time-to-resolution
- Feedback loops: learning from operator overrides and outcomes
- Tracking costs and managing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled jobs
- Implementing gradual autonomy rollout: from shadow mode to full auto-remediation
- Handling agent failures: protocols for when the agent itself malfunctions
- Building the business case and measuring ROI for autonomous operations
Requirements
- Hands-on experience with IT operations, DevOps, or SRE practices.
- Familiarity with Python scripting and REST APIs.
- A foundational understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI for incident management.
14 Hours