Infrastructure • Automation • Cloud • Innovation

Kranthi Veeravalli

Senior Infrastructure & Automation Engineer

I engineer reliable enterprise infrastructure, automate complex operations, modernize platforms, and turn practical ideas into working solutions.

13+ YearsEnterprise experience
EnterpriseModernization & migrations
HybridCloud & infrastructure
AutomationAt enterprise scale
Engineer
Build • Automate
Modernize • Innovate
VMware
Nutanix
Cloud
Storage & SAN
Ansible
What I Engineer

Enterprise Capabilities

Not a wall of logos. Each capability connects to real infrastructure responsibilities, migrations, automation and operational outcomes.

Infrastructure & Virtualization

VMware vSphere, vCenter, SRM, VCF, Nutanix AHV, Windows and Linux infrastructure.

Storage, SAN & Data Protection

NetApp, Hitachi VSP, Infinidat, Dell EMC VMAX/VNX/PowerMax/PowerStore/Isilon, Cisco MDS, Brocade, replication and data protection.

Automation & Operations

Ansible, AWX/AAP, PowerShell, REST APIs, CI/CD, ServiceNow and large-scale infrastructure automation.

Cloud & Modern Platforms

Azure, Azure Local, Azure Arc, GCP, Kubernetes, monitoring and AI/GPU infrastructure.

Selected Engineering Impact

Technology → Measurable Results

Delivered outcomes alongside forward-looking AI + infrastructure initiatives currently being explored.

~3,000 VMs

VMware → Nutanix Migration

Direct migration scope across production, test and sandbox, including VDI and application workloads.

495 SAN Switches

Ansible Code Upgrades

Cisco SAN upgrade automation across a two-location environment; broader upgrade cycle improved from roughly six months to about one month.

46 Arrays

Enterprise Storage & DR

Storage/SAN operations, replication, DR and data protection across enterprise platforms.

AI/ML Compute

Dell R7725 + NVIDIA GPUs

GPU server build and platform readiness for ornithology AI/ML workloads, including PCIe, BIOS/iDRAC, thermal and infrastructure considerations.

Observability

Prometheus + Grafana

GPU utilization, storage capacity and compute-health dashboards for proactive infrastructure monitoring.

ITSM + Logs

Splunk + ServiceNow

Operational pattern for querying infrastructure logs, correlating alerts/incidents and routing actionable events into ITSM workflows.

Hybrid Modernization

VMware → Azure Local

Workload assessment, VLAN/logical-network mapping, Azure Migrate replication, phased cutover, validation and rollback planning.

In Progress / Exploration

Vertex AI + Gemini

Proactive Infrastructure Intelligence

Exploring telemetry-driven analysis for anomaly, capacity and performance-risk insight across hybrid infrastructure.

In Progress / Exploration

NVIDIA NIM + Enterprise LLM

Continuous Log, Incident & RCA Intelligence

Exploring continuous log/event correlation, incident-draft assistance and RCA acceleration using enterprise LLMs with engineer review.

Engineering Stories

From Challenge to Working Solution

Click a story to see the challenge, engineering approach and outcome.

01 — Enterprise Modernization: VMware → Nutanix AHV

Environment: ~10,000+ VM enterprise estate across production, test and sandbox.

My scope: ~3,000 VMs, including VDI and application workloads, including systems supporting Oracle databases.

Platform: Nutanix clusters built on Dell server infrastructure.

02 — Automation at Scale: 495 Cisco SAN Switches

Environment: 46 enterprise storage arrays and 495 Cisco SAN switches distributed across two locations, supporting production, test and sandbox environments.

The challenge: SAN switch software upgrades were performed manually. Engineers had to stage the appropriate Cisco software images on each switch, validate available device storage and compatibility, and execute the upgrade workflow. At fleet scale, a complete upgrade cycle could take approximately six months.

Automation approach: We developed Ansible playbooks for Cisco SAN switch upgrades and initially validated the workflow against test switches before expanding the rollout. Upgrade jobs were launched through the enterprise Ansible controller (AWX / Ansible Tower, depending on the platform in use), with the playbooks maintained as controller projects/job templates.

Operational improvement: A switch upgrade could typically complete in roughly 20 minutes, depending on chassis and module count. With controlled automation and parallelized rollout planning, the broader switch upgrade cycle was reduced from roughly six months to about one month.

Challenges encountered: Automation was not simply “write a playbook and run it.” We encountered controller job failures and suspended jobs, controller resource/capacity constraints, insufficient switch filesystem/bootflash space for software images, upgrade tasks terminating mid-workflow, and timeout issues during long-running device operations.

Engineering lessons: The solution evolved to emphasize pre-checks before change execution: image and platform validation, available-space checks, controller capacity, timeout tuning, controlled batches, failure handling, post-upgrade validation and safe recovery procedures.

Why this matters: This case study demonstrates automation at infrastructure scale, including the operational failures and safeguards required to turn a manual network-maintenance process into a repeatable production workflow.

03 — Resilience: VMware SRM & Storage Replication

Challenge: Improve recovery repeatability and reduce DR failover effort.

Engineering: SRM protection groups, automated recovery plans, RTO/RPO validation and storage replication architecture.

04 — Modern Infrastructure: GPU, AI/ML & Observability

Research use case: The Lab works with bird-observation data captured through distributed camera systems. The resulting data can support research into bird activity, behavior and movement patterns and can also contribute to public-facing educational content.

My infrastructure scope: Build and configure Dell PowerEdge R7725 GPU server infrastructure according to the compute, memory, GPU, storage, networking and availability requirements of AI/ML research workloads. My role is infrastructure enablement—not development of the bird-analysis AI/ML models.

Compute platform: The PowerEdge R7725 is a 2U dual-socket AMD EPYC 9005 platform supporting up to 192 CPU cores per socket. Depending on the validated configuration, it can support accelerator options including NVIDIA L4, L40S, H100 NVL, H200 NVL and other supported GPUs.

From Challenge to Working Solution

Infrastructure Integrations That Connect the Environment

Not every engineering win is a migration. Some of the most valuable work is connecting platforms so operations become safer, faster and more observable.

Certificate Lifecycle Automation

Splunk/CMDB expiry visibility → ServiceNow ticket → PowerShell CSR → Microsoft CA → Azure Key Vault → approval → Ansible deployment/activation across Windows, Linux, vCenter, storage, Cisco and load-balancer/API infrastructure.

Backup & Recovery Integration

On-prem infrastructure protection with enterprise backup/recovery, cloud-aware architecture, replication and validation so recovery is designed into modernization rather than added afterward.

Logs → Incident Operations

Splunk queries and infrastructure events → investigation/correlation → ServiceNow incident/change/problem workflow → RCA and governed remediation.

Monitoring Integration

Prometheus/Grafana infrastructure dashboards plus Splunk operational investigation for compute, GPU, storage and platform health.

Storage Replication & DR

Replication between enterprise arrays, VMware SRM protection/recovery planning, failover validation and RTO/RPO-focused operations.

Automation Across Infrastructure

Ansible/AWX-AAP and PowerShell applied beyond one device type: SAN upgrades, health checks, certificates and repeatable infrastructure operations.

Lessons Learned

What Real Infrastructure Delivery Taught Me

Successful infrastructure work is not only the final architecture. It is planning, implementation, failures, recovery, validation and the lessons carried into the next change.

N

Nutanix Platform Transition

Cluster setup on Dell servers, migration readiness, workload transition, operational validation and the challenges encountered while moving enterprise workloads from VMware to Nutanix.

Nutanix Lifecycle & Upgrade Thinking

Upgrade planning should begin with compatibility, health and dependency checks—not with the upgrade button. Controlled sequencing and recovery readiness are part of the design.

A

Automation Failure Is Engineering Data

AWX failures, timeouts, resource constraints and device-side storage issues shaped better pre-checks, batching, retry and validation strategies.

Successful Transition Requires Validation

Migration success is more than moving a VM: application reachability, storage, network, database dependencies, performance and operational handoff must all be validated.

Kranthi's Innovation Lab

Beyond the Enterprise

Personal projects. Real problems. Working solutions. A place to show curiosity, product thinking and practical cloud engineering.

FLAGSHIP PROJECT

BhoomiSetu

Andhra Pradesh Land Intelligence Platform

Land information, marketplace intelligence, location validation, agriculture insights and a bilingual English/తెలుగు experience.

DataGCPMarketplaceVerificationEnglish | తెలుగు

Deal Guru

Retail Value Intelligence

Compare products across retailers by matching quantities and normalizing unit price—not just sticker price.

APIsData ProcessingCloud

Advaitha

School Assistant

A cloud-based reminder assistant using Telegram, Cloud Run, Firestore and scheduled notifications.

TelegramCloud RunFirestore
AI + Infrastructure • Exploration Vertex AI + Gemini for proactive infrastructure analysis  •  Nutanix Enterprise AI + enterprise LLMs for logs, incidents and RCA assistance.
Career Journey

Experience That Evolved With Technology

A concise recruiter view. The final site can expand each role into selected deliveries without reproducing the entire résumé.

March 2026 — Present

Cornell University — Lab of Ornithology

Senior Lead Infrastructure Administrator

VMware and Azure Local modernization, AI/GPU infrastructure, Kubernetes, storage/backup, Active Directory migration and platform operations.

March 2023 — December 2025

HCL Technologies

Senior Lead Infrastructure Engineer

Large-scale Nutanix migrations, VMware, enterprise storage, SAN, automation, cloud storage integration and observability.

May 2021 — January 2023

NTT Data

Systems Integration Advisor

Hybrid infrastructure integration, storage migrations, NetApp, Dell EMC, Nutanix and enterprise project delivery.

2012 — 2021

Mphasis • DXC Technologies

Infrastructure Engineering & Leadership

SAN/NAS engineering, enterprise storage migrations, automation concepts, operations leadership and large end-of-life modernization programs.

Recruiter View

One Link. Multiple Depths.

A recruiter can understand the profile quickly; a hiring manager can open engineering stories; a technical interviewer can explore architecture and design decisions.

30 Seconds

Professional identity, core capabilities and measurable impact.

3 Minutes

Career journey, selected engineering deliveries and Innovation Lab.

15 Minutes

Architecture, technical decisions, challenges, troubleshooting and lessons learned.

Resume

The conventional document remains available as a supporting artifact—not the entire experience.

View Resume → Download DOCX →

Prototype note: public launch should remove phone details from the page, verify all quantified claims, confirm employer-confidentiality boundaries, and connect final Resume / LinkedIn / project URLs.

Email copied