Infrastructure Leader
DevOps | Hybrid in Islamabad, Punjab Province and Federal Capital Territory (ICT), Pakistan | Full Time
Job Summary
CONVO is seeking a Senior Infrastructure Leader to own the Kubernetes and cloud foundation for its agentic platform. You’ll lead the architecture and operations of scalable, secure, and resilient cloud infrastructure across environments, with a strong focus on Kubernetes, Azure, Terraform, GitOps, observability, identity, workload isolation, and cost optimization. You’ll also support reliable LLM-serving workloads and establish infrastructure standards, capacity controls, and operational practices for the platform.
Technical mission
Own the Kubernetes and cloud foundation for the agentic platform, including scalability, isolation, observability, identity integration and cost-per-request control.
Key responsibilities
- Define the target Kubernetes and cloud topology for development, test and production environments.
- Own autoscaling, workload placement, network boundaries and runtime isolation for agent, tool and LLM-serving workloads.
- Set infrastructure-as-code, GitOps, identity, secrets, gateway and certificate-management standards.
- Define service-level indicators, operational dashboards, alerting and incident-response expectations for the platform.
- Establish capacity and FinOps controls, including cost-per-request budgets and resource-consumption attribution.
- Review platform changes for security, resilience, recoverability and vendor decoupling.
Required technical capabilities
- Minimum experience: 8+ years in cloud/platform infrastructure, including 3+ years leading production Kubernetes architecture or operations.
- Production architecture and operations experience with Kubernetes, including autoscaling, networking, storage and workload isolation.
- Strong Terraform or equivalent infrastructure-as-code capability and practical GitOps delivery experience.
- Hands-on cloud platform experience, preferably Azure, covering compute, networking, managed identity and secure service exposure.
- Observability engineering across metrics, logs and traces, with actionable service and cost dashboards.
- Experience integrating identity providers, API gateways, secrets management and policy controls.
- Understanding of LLM inference or model-serving workloads, including GPU/CPU scheduling and endpoint reliability.
Preferred experience
- AKS, Azure Monitor/Application Insights, managed identity and private networking.
- Policy-as-code, service mesh, multi-cluster operations and disaster-recovery design.
- FinOps practices for shared AI platforms and usage-based cost allocation.
Expected deliverables / acceptance evidence
- Approved cloud/Kubernetes reference architecture and environment topology.
- Version-controlled IaC and GitOps baseline with security and rollback controls.
- Capacity, availability and cost-per-request dashboards with alert thresholds.
- Operational runbooks for deployment, failure recovery and platform incidents.
Primary interfaces
- Works with the Agent Runtime team, DevOps, Architecture Office, security/identity owners and the data-platform infrastructure lead.
Why Join CONVO?
At CONVO, we’re reimagining how sales organizations in FMCG run smarter, faster, and more profitably. You’ll be joining a team of thinkers, builders, and consultants who thrive at the intersection of business and technology.
We offer:
- Global exposure working with Tier-1 FMCG clients
- A high-impact role where your input shapes real-world commercial outcomes
- A collaborative culture with mentorship, continuous learning, and career growth opportunities
