Semua lowongan
Logo Grab

Lead Infrastructure Engineer, Managed Kubernetes Platform (MEKS)

Grab

LokasiJakarta, Indonesia
NegaraIndonesia
StatusFull-time
Diposting15 Agustus 2026

Deskripsi pekerjaan

Get to Know the Role

The Lead Infrastructure Engineer – Kubernetes & Service Mesh is a lead technical authority within the MEKS platform team. This role is responsible for designing, building, and operating the high-scale container and service mesh infrastructure. This infrastructure powers critical backend services across Grab. Reporting directly to the Platform Engineering Manager, this hands-on lead role focuses entirely on deep technical execution, platform architecture, system reliability, and developer experience. You will serve as the primary architect and technical mentor for the platform, driving multi-cluster AWS EKS strategies, Istio service mesh implementations, and cross-team developer enablement. This is an onsite position based in the Jakarta office.

The Critical Tasks You Will Perform** **Architecture & Platform Engineering

  • Service Mesh Implementation involves designing and deploying production-grade Istio service mesh infrastructure. This infrastructure consists of a data plane and control plane. It requires managing various aspects, including mTLS, traffic management, canary/blue-green deployments, rate limiting, and circuit breaking, to optimize the service mesh.
  • Architect and manage large-scale AWS EKS clusters, driving multi-cluster/multi-region topologies, node-pool optimization, custom scheduler strategies, and automated cluster lifecycles.
  • Establish standardized cluster templates, namespaces, network policies, RBAC, and policy-as-code guardrails across the organization.

Developer Experience & Self-Service Tooling

  • Build and maintain self-service APIs, GitOps workflows, and automation that simplify onboarding, deployments, and traffic management for internal product engineering teams.
  • Partner directly with application teams to guide architectural choices, streamline workload migrations to MEKS, and accelerate time-to-market.
  • Create comprehensive platform documentation, golden path templates, runbooks, and reference architectures to foster engineering self-sufficiency.

Reliability, Observability & Incident Response

  • Implement telemetry, metrics, distributed tracing (OpenTelemetry/Jaeger), and centralized logging across clusters, Envoy proxies, and application workloads.
  • Establish and uphold technical SLOs/SLAs, driving proactive capacity management, resource right-sizing, and performance tuning for peak availability and optimal infrastructure efficiency.
  • Act as a senior technical point of escalation for complex platform outages, lead root-cause analyses , and execute corrective actions to prevent recurrence.

Technical Mentorship & Collaboration

  • Mentor mid-level and senior engineers on the team in modern Kubernetes, mesh architecture, and distributed systems best practices.
  • Collaborate with Security, SRE, Networking, and Cloud Infrastructure teams to ensure platform security, regulatory compliance, and seamless integration.

Tertarik dengan posisi ini?

Lamaran diproses di situs grab.careers — Kerjago tidak menyimpan lamaran untuk lowongan ini.