Essential Duties and Responsibilities
-
Designs, implements, and manages AWS cloud infrastructure, including VPCs, EC2, RDS, S3, IAM, Lambda, and other foundational services.
-
Develops and maintains Terraform modules to provision and manage infrastructure in a repeatable, version-controlled manner.
-
Deploys, manages, and optimizes Kubernetes clusters (EKS), handling workload orchestration, autoscaling, and resource management.
-
Builds and improves continuous integration and deployment (CI/CD) pipelines to enable fast, safe, and automated releases.
-
Implements logging, monitoring, and alerting solutions to ensure system health and enable rapid incident response.
-
Applies strict infrastructure security best practices, manages secrets, and ensures compliance with internal security policies.
-
Writes scripts and tools to automate operational tasks, reducing manual toil and improving the overall developer experience.
-
Creates and maintains runbooks, architecture diagrams, and technical documentation.
-
Partners with development teams to improve platform reliability and streamline application deployments.
-
Participates in an on-call rotation schedule to respond to production incidents and ensure continuous system reliability.
-
Designs and provisions cloud infrastructure specifically tailored to support scalable AI/LLM workloads and agentic systems, ensuring optimized compute resource allocation and high availability.
-
Evaluates and integrates AI-driven DevOps tools and LLM-assisted automations into the platform workflow to accelerate incident response, infrastructure provisioning, and operational troubleshooting.
Qualifications
-
Location: This role is fully remote and open exclusively to candidates based in LATAM.
-
Experience:
-
4+ years of proven experience in infrastructure, DevOps, or platform engineering.
-
Strong proficiency with AWS services and modern cloud architecture patterns.
-
Hands-on experience utilizing Terraform for Infrastructure as Code (IaC).
-
Solid understanding of Kubernetes concepts (deployments, services, ingress, RBAC, Helm charts) and containerization utilizing Docker.
-
Proficiency in TypeScript and Bash scripting.
-
Familiarity with standard CI/CD tools (e.g., GitHub Actions, GitLab CI, ArgoCD).
-
Strong understanding of networking fundamentals (DNS, load balancing, firewalls, VPNs).
-
Hands-on experience provisioning infrastructure for AI agent frameworks (e.g., LangChain, LlamaIndex, CrewAI) or deploying LLMs in a production or near-production context.
-
Strong troubleshooting, problem-solving, and cross-functional communication skills.
-
-
Nice to Have:
-
Experience coding in Golang.
-
Knowledge of GitOps practices and tools (ArgoCD, Flux).
-
Familiarity with modern observability stacks (Prometheus, Grafana, Datadog, ELK).
-
AWS certifications (Solutions Architect, DevOps Engineer) or Kubernetes certifications (CKA/CKAD).
-
Experience with service mesh implementations (Istio, Linkerd) and FinOps practices for cost optimization.
-
