- Own end-to-end delivery of ERP migration, pricing analytics, retail sales integration, and product information initiatives on an Azure Databricks Lakehouse (medallion / Delta Lake) architecture.
- Led migration of 180+ ERP source tables to Delta Lake, building Python DAGs orchestrated in Apache Airflow, replacing legacy on-prem batch processing with zero schema-change handling errors.
- Designed 30+ production Airflow pipelines covering 180+ tables, raising pipeline reliability to 99.999% SLA.
- Engineered ingestion and modeling of Walmart retail sales data at 1.8B+ row scale, tuning partitioning, Z-ordering, and liquid clustering, and leveraging the Photon accelerator to speed up Spark query execution for 10× faster analytical queries.
- Authored and tuned T-SQL stored procedures, functions, and triggers in Azure SQL DB for ERP validation logic, applying performance tuning and query optimization.
- Configured Always On Availability Groups, replication, and clustering (WSFC) for Azure SQL DB high availability and disaster recovery.
- Implemented Databricks Asset Bundles (DABs) for CI/CD-deployed, versioned pipeline releases; enabled cross-org data sharing via Delta Sharing.
- Extended the Databricks lakehouse into Microsoft Fabric / OneLake to support enterprise reporting workflows for downstream BI consumers.
- Mentored engineering management and the broader team on the lakehouse architecture end-to-end — from design decisions through production implementation — building shared technical ownership across the organization.
- Built and deployed an AI code-review agent using Claude Code, applying MCP (Model Context Protocol) tool orchestration and agentic engineering to automate pull-request review directly in the CI/CD pipeline.
Sankar Sundaram
10+ years as a Data/AI engineer and Cloud/AI solutions architect designing and deploying enterprise-scale data platforms, with recent hands-on focus on AI/LLM engineering and agentic workflows — building large-scale ELT/ETL and lakehouse platforms on Databricks, API ingestion, Delta Sharing, and orchestration across Airflow, dbt, and Python. Depth in enterprise data warehousing: dimensional modeling, T-SQL, stored procedures, and high-availability SQL Server, with a proven record migrating legacy on-prem data warehouses to modern Delta Lake and optimizing pipeline reliability to 99.999% SLA. Adept at driving CI/CD initiatives and collaborating across teams to deliver secure, scalable, AI-powered solutions — including LLM-integrated agents and MCP-based tool orchestration built directly into the delivery pipeline.
603-930-7755 [email protected] Jacksonville, FL LinkedIn ↗ GitHub ↗- 10+
- Years in data engineering & cloud architecture
- 1.8B+
- Rows ingested in a single retail sales pipeline
- 99.999%
- SLA reliability across 30+ production pipelines
- 180+
- ERP source tables migrated to Delta Lake
Core expertise
Cloud & Lakehouse Platforms
- Azure Databricks & Delta Lake
- Google BigQuery & GCP
- Azure Data Factory, Synapse & Microsoft Fabric
- Azure SQL DB & SQL Server
Orchestration & Transformation
- Apache Airflow (Python DAGs)
- dbt (Data Build Tool)
- Spark SQL & PySpark
- Kafka & Spark Structured Streaming
- Databricks Asset Bundles
Data Modeling & SQL Engineering
- T-SQL, Stored Procs, Functions & Triggers
- Performance Tuning & Query Optimization
- Data Warehousing — Dimensional, Medallion & Data Vault Modeling
- Unity Catalog & Governance (RBAC, Lineage)
Delivery & Platform Ops
- CI/CD — Azure DevOps, GitHub, Bitbucket
- Terraform, Docker & Kubernetes
- Power BI, Tableau, SSRS & SSIS
- Agile / Scrum & SDLC
Agentic Engineering & GenAI
- Claude Code, Cursor & Codex
- MCP (Model Context Protocol)
- Agent Frameworks & Tool Orchestration
- Python Development
- Generative AI & LLM Integration
Experience
- Migrated an on-prem SQL Server enterprise data warehouse to Google BigQuery, reducing report-refresh latency from 20 minutes to 2 minutes.
- Migrated legacy T-SQL stored procedures, functions, and triggers into BigQuery SQL and Databricks, applying database design and performance-tuning best practices.
- Developed dbt transformation pipelines structuring raw data into curated gold-layer models for self-service BI.
- Built Cloud Run services ingesting API data into BigQuery and Snowflake; established DevOps best practices in Azure DevOps.
- Applied agentic engineering and Python-based GenAI/LLM integration to prototype workflow-automation agents, evaluating and using multiple AI coding tools and LLMs — including Claude Code, Cursor, and Codex — and comparing agent frameworks and tool-orchestration patterns for practical business use cases.
- Designed large-scale big-data pipelines for sustainability analytics on Azure Databricks and ADF/Synapse, improving reliability to 99.999% SLA.
- Built ADF/Databricks pipelines across ADLS Gen2, Azure SQL, Event Hub, and Logic Apps; implemented governance and lineage with Microsoft Purview.
- Delivered interactive Power BI and Tableau dashboards using PySpark, Databricks notebooks, and Spark SQL.
- Delivered big-data migration and API integrations on GKE, Cloud Dataflow, Cloud Run, and Cloud SQL; built an analytics platform on Snowflake and SnowSQL.
- Ingested data via AWS Glue, S3, and Lambda; ran CI/CD with Kubernetes, Docker, Jenkins, and Bitbucket.
- Built API integrations with Veeva Vault CRM and Veeva Link to ingest field-sales and medical-affairs data into the Snowflake analytics platform for commercial reporting.
- Supported regulatory and quality data flows by integrating pipelines with Veeva RIM and Veeva QMS, ensuring traceable, audit-ready data for compliance reporting.
- Led Veeva LIMS implementation end-to-end — requirements, workflow design, configuration, integration, and testing — partnering with Lab, Quality, IT, and Business Process Owners through UAT and go-live.
- Applied risk-based CSV/CSA and GxP principles across Veeva LIMS configuration and documentation, aligning with 21 CFR Part 11, EU Annex 11, GAMP 5, and ALCOA+.
- Built ETL/ELT and ML pipelines using ADF, Azure Databricks, Spark, and Python for R&D and IoT/sensor streaming data.
- Designed star/snowflake schema models in Snowflake; automated IaC deployment with Terraform and Azure DevOps.
Education & certifications
Master of Science in Cyber Security (MSCS)
EC-Council University, Albuquerque, NM · GPA 3.80 / 4.0
Certifications
- Google Certified Professional Cloud Architect
- Microsoft Certified: Azure AI Engineer Associate
- Databricks Lakehouse Fundamentals
- Certified Gen AI for Data Engineers
- Google Cloud Digital Leader
Earlier experience
Migrated SAS reporting to SQL/Azure Data Lake; built ETL (SSIS, ADF) and Power BI/Tableau analytics.
Led enterprise systems integration; built Azure Data Lake, Informatica, and Oracle PL/SQL data ingestion.
Delivered BI/DW solutions (Informatica, OWB, Tableau, SAS) and MDM for life-sciences applications, integrating feeds from Veeva Vault (QualityDocs, eTMF).