Professional Profile
Results-driven Cloud and DevOps Engineer with a strong background in Linux systems
administration, cloud infrastructure, and automation: born on-prem and forged in the cloud.
Proven experience managing Kubernetes clusters, deploying cloud solutions, automating
infrastructure, and optimizing system performance. Passionate about leveraging open-source
technologies to improve scalability, reliability, and operational efficiency across modern
cloud environments.
Technical Skills
- AWS
- Kubernetes
- Terraform
- Linux
- Docker
- Azure DevOps
- GitHub Actions
- GitLab
- Helm
- Argo CD
- Ansible
- Bash
- Python
- Prometheus
- Grafana
- Kibana
- Datadog
- OpenStack
Professional Experience
-
Design, develop, and maintain CI/CD pipelines using Azure DevOps to support automated
application delivery and infrastructure workflows.
-
Collaborate with development and infrastructure teams to streamline build,
testing, and release workflows.
-
Participate in the migration of CI/CD workflows from Azure DevOps to GitHub Actions,
contributing to pipeline modernization efforts.
-
Troubleshoot and resolve pipeline, deployment, and automation-related issues across
development and production environments.
-
Supported AWS infrastructure in multi-region Terraform-managed environments,
ensuring service availability and resilience across regions.
-
Coordinated deployments across multiple teams and services, validating
dependencies and minimizing service disruption during releases.
-
Deployed, configured, and troubleshot Amazon EKS Kubernetes clusters,
including node scaling, networking issues, and workload failures.
-
Diagnosed and resolved AWS IAM permission and access issues, including
policy creation, least-privilege adjustments, and cross-service access
debugging.
-
Used Amazon CloudWatch for monitoring, alerting, and infrastructure
diagnostics.
-
Performed troubleshooting using Datadog metrics, log streams,
dashboards, and alerting to accelerate incident resolution.
-
Followed and improved operational runbooks, documentation,
and standard operating procedures.
-
Managed AWS cloud services including RDS, S3, Lambda,
Bedrock, and SageMaker.
Key Achievements
-
Improved infrastructure reliability by standardizing Terraform modules
and deployment patterns, reducing deployment-related incidents by
approximately 25%.
-
Contributed to reducing MTTR by approximately 30% through enhanced
Datadog and CloudWatch observability.
-
Improved shared Terraform modules, increasing reusability,
operational consistency, and deployment efficiency.
-
Migrated legacy infrastructure into modular Terraform structures,
reducing onboarding time for new engineers.
-
Collaborated in coordinated production deployments, reducing rollout
friction and minimizing downtime.
-
Supported self-managed Kubernetes platforms deployed with Kubespray,
including cluster operations, workload troubleshooting,
Helm deployments, and GitOps workflows with Argo CD.
-
Investigated infrastructure, networking, and platform-related issues
across hybrid and public cloud customer environments.
-
Managed observability platforms using Prometheus, Grafana,
and Kibana, including dashboards, alert tuning,
and centralized log analysis.
-
Administered Longhorn distributed storage,
persistent volumes, backup and restore operations,
and storage troubleshooting.
-
Automated operational tasks using Bash,
Python, and Ansible.
Key Achievements
-
Promoted to a higher-responsibility role based on performance.
-
Ranked among the Top 3 performers in ticket resolution during Q4 2024.
-
Reduced incident resolution times by improving monitoring,
alerting, and proactive issue detection.
-
Performed first-level troubleshooting and incident triage for
infrastructure, networking, Kubernetes, and OpenStack environments.
-
Investigated alerts by analyzing logs, metrics, and monitoring data
before escalating incidents to engineering teams.
-
Managed incident tracking and customer communication through Jira
ticketing workflows.
-
Supported Kubernetes and OpenStack-based environments across
hybrid and multi-datacenter infrastructures.
Key Achievements
-
Improved incident response efficiency through proactive monitoring
and faster alert triage.
-
Reduced unnecessary escalations by resolving infrastructure and
platform issues at the first support level.
-
Developed strong troubleshooting expertise across Linux,
Kubernetes, networking, and cloud infrastructure.
-
Supported AWS cloud infrastructure with emphasis on EC2 lifecycle
management, Linux administration, and operational excellence.
-
Coordinated maintenance windows and Linux server patching across
production environments.
-
Enhanced infrastructure observability through Datadog integrations
for monitoring, log aggregation, and troubleshooting.
-
Supported automation initiatives using GitHub Actions for routine
operational workflows.
Key Achievements
-
Supported more than 100 Linux servers running on AWS.
-
Maintained 99.9% availability for critical production workloads.
-
Reduced manual operational work through process standardization and
monitoring improvements.
-
Assisted patching operations across multiple Linux environments with
minimal downtime.
-
Led automation efforts by creating reusable Bash, Python,
and Ansible solutions adopted across operational teams.
Education
Computer Engineer
Faculty of Mathematics
Universidad Autónoma de Yucatán
2018
Certifications
-
Google IT Support Professional Certificate — Coursera (2023)
-
DevOps Engineer Bootcamp — ITJuana (2023)