Este perfil está escrito en inglés
Demonstrated experience with Apache Spark.
Demonstrated experience with Kafka.
Demonstrated experience with Hive.
Demonstrated experience with Data Lake.
Demonstrated experience with Apache Hadoop.
Professional profile
Designed and optimized scalable data pipelines using Azure Databricks and PySpark, improving processing efficiency and supporting high-volume data workloads.. Developed batch and near-real-time ETL workflows using Azure Data Factory and Azure Synapse Analytics for enterprise data integration and transformation.. Implemented Delta Lake architecture to improve data reliability, consistency, and transactional processing across analytical datasets.. Built and maintained Spark-based data processing applications for large-scale transformation and distributed workloads.. Developed and managed Apache Airflow workflows to orchestrate, schedule, and monitor multiple data pipelines.. Designed and optimized data solutions using SQL Server and Azure SQL to support high-throughput applications and analytical workloads.. Automated build and deployment processes using Azure DevOps and Jenkins, improving consistency across development and production environments.. Implemented data governance, security, and compliance practices across cloud-based data platforms.. Supported reporting and analytics initiatives by preparing curated datasets for Power BI and Tableau dashboards.. Collaborated with cross-functional teams to develop and support streaming data solutions for business and operational requirements.. Modernized legacy ETL processes by migrating workloads to scalable cloud-based architectures.. Designed data lake storage and processing strategies to support efficient ingestion, transformation, and retrieval of enterprise data.. Enhanced metadata management and data lineage capabilities using Informatica and Control-M.. Developed reusable Python and SQL scripts to automate data transformations, validation, and operational processes.. Supported machine learning workloads within Azure Databricks by preparing and processing datasets for predictive analytics.. Integrated Kafka and Azure Event Hubs to support real-time and event-driven data ingestion.. Implemented data partitioning and optimization strategies in Databricks to improve query and processing performance.. Developed mechanisms to manage schema changes for structured and semi-structured datasets.. Implemented role-based access controls (RBAC) for Azure Data Lake and Databricks environments.. Improved pipeline reliability through enhanced observability, logging, monitoring, and operational support.. Optimized Databricks cluster utilization and lifecycle management to improve resource efficiency.. Developed scalable ETL processes to support high-volume and continuously generated data.. Tuned Spark applications by optimizing memory usage, compute resources, and job execution.. Implemented automated recovery and retry mechanisms to improve data pipeline resiliency.. Developed automated data quality validation using Great Expectations and PySpark.. Supported distributed machine learning and large-scale data preparation workflows within Databricks.. Developed and maintained data catalog and metadata capabilities using Azure Purview.. Built serverless data processing components using Azure Functions for event-driven and automated workloads.. Implemented data masking approaches to protect sensitive information stored in Azure SQL and Synapse.. Developed containerized data processing workloads using Docker and Kubernetes.. Created Python-based integration components to connect third-party APIs with Databricks data workflows.. Implemented serverless and event-driven processing using Azure Functions and Event Grid.. Developed automated metadata management processes for datasets stored in Azure Data Lake.. Implemented monitoring and alerting solutions using Azure Monitor and Log Analytics to improve visibility into pipeline performance and failures.
Designed and supported enterprise data warehouse solutions using Snowflake for scalable analytics and reporting.. Developed ETL and ELT pipelines using Azure Data Factory to move and transform data across cloud and enterprise systems.. Designed and optimized data lake architectures using Azure Data Lake Storage and Azure Synapse Analytics.. Developed complex SQL solutions for data transformation, validation, and processing within Snowflake and Azure SQL.. Built and managed automated data ingestion workflows using Informatica Scheduler.. Implemented pipeline monitoring, scheduling, and alerting processes using AutoSys and Control-M.. Managed source code and version control practices using Git and Bitbucket for data engineering projects.. Supported analytics and reporting initiatives by preparing datasets for Power BI and Tableau.. Strengthened data governance and compliance processes through automated validation and security controls.. Developed Python-based data processing and cleansing solutions for enterprise datasets.. Supported migration of on-premises SQL databases and workloads to Azure SQL and Snowflake.. Implemented RBAC and data security practices to manage user access across cloud data platforms.. Developed and maintained processes supporting data lineage and metadata management.. Built distributed data processing solutions using Apache Spark and Hadoop.. Automated data backup and recovery processes to support data availability and business continuity requirements.. Developed serverless data processing workflows using Azure Functions.. Improved Snowflake query performance through clustering strategies, query optimization, and warehouse management.. Developed Python-based connectors and integrations for third-party systems and Snowflake data workflows.. Supported migration of legacy Hadoop-based workloads to cloud platforms including Azure Synapse.. Optimized Snowflake warehouse utilization and workload management to improve resource efficiency.. Developed streaming and near-real-time data pipelines using Azure Stream Analytics.. Supported data integration workflows across AWS and Azure cloud environments.. Implemented data retention and storage management strategies within Snowflake.. Automated user provisioning and access management processes for Azure SQL and Snowflake.. Enhanced ETL error handling through structured logging, monitoring, and exception management.. Supported event-driven data ingestion architectures for high-volume device and operational data.. Created and optimized Snowflake materialized views to improve query performance for analytical workloads.. Developed data models using Azure Cosmos DB to support specialized data processing requirements.. Supported AI and recommendation-based analytics initiatives using curated datasets stored in Snowflake.. Integrated Kafka and Azure Event Hubs for high-throughput and streaming data ingestion.. Supported feature engineering and data preparation workflows for advanced analytics use cases.. Developed containerized data processing applications using Docker and Kubernetes.. Implemented automated retry and recovery strategies to improve reliability of critical ETL workflows.. Participated in architecture reviews and contributed to data engineering standards and best practices.. Integrated Azure Data Factory with multiple cloud storage platforms and enterprise data sources.
Designed and developed cloud-based big data solutions using AWS Glue, EMR, and Amazon Redshift.. Built scalable data processing pipelines using Apache Spark and Databricks.. Developed ETL workflows using AWS Glue and Amazon Athena for large-scale data processing and analytics.. Supported migration of on-premises databases to Amazon RDS and Amazon Redshift.. Developed SQL-based transformation logic for processing high-volume enterprise datasets.. Designed and managed cloud data lake architectures using Amazon S3 and Delta Lake.. Implemented data governance and security controls using AWS IAM and AWS DMS.. Prepared curated datasets to support reporting and visualization requirements in Power BI and Tableau.. Automated build and deployment processes using Bitbucket, AWS CodeCommit, and Jenkins.. Designed batch and near-real-time data processing workflows to support business reporting and operational requirements.. Supported performance tuning and optimization of Hadoop-based data processing environments.. Developed Shell scripts to automate recurring data processing and operational activities.. Created and maintained Apache Airflow DAGs for scheduling, dependency management, and workflow orchestration.. Integrated multiple enterprise data sources into centralized cloud-based data platforms.. Implemented cloud resource and processing optimizations to improve infrastructure utilization.. Developed Redshift Spectrum queries to enable analysis of data stored within Amazon S3.. Built event-driven data processing workflows using AWS Lambda and Amazon Kinesis.. Developed serverless data processing solutions to support scalable and automated workloads.. Implemented secure data-sharing processes using AWS cloud data-sharing capabilities.. Designed and supported real-time streaming pipelines using Kafka and Amazon MSK.. Configured CloudWatch and CloudTrail for monitoring, logging, and operational visibility.. Supported multi-region data replication and disaster recovery strategies.. Implemented metadata and data catalog capabilities using AWS Glue Data Catalog.. Developed and managed S3 lifecycle policies to support data retention and storage optimization.. Created Python and PySpark applications for custom data transformation and processing requirements.. Performed Spark performance tuning to improve execution efficiency and resource utilization.. Collaborated with business and technical teams to improve data architecture and processing strategies.. Implemented schema evolution strategies for data pipelines built using AWS Glue and Databricks.. Designed and optimized Redshift table structures and query strategies to improve analytical performance.. Developed data quality and anomaly detection processes to identify potential issues within data pipelines.. Created log analysis and monitoring workflows using the ELK Stack (Elasticsearch, Logstash, and Kibana).. Provided technical guidance and mentoring to junior engineers on big data development practices and data engineering standards.