São Paulo, Brazil

Hi, I'm Victor. Data Engineer.

Data Engineer with 5+ years of experience building robust pipelines in cloud architectures. Specialized in legacy system migration, ETL/ELT processing with Spark, and data infrastructure on AWS/GCP. Focused on data quality, integrity and availability using open source technologies like Spark, Kafka and Airflow.

Experience

  1. 2025 — Now

    Data Specialist · Serasa Experian (via ACT Digital)

    Led a complex mainframe (Cobol/DB2) to AWS pipeline migration processing massive daily positional file ingestion. Implemented data validation and quality with consolidated table generation (bronze/silver/gold) for ML and other products. Production deployment with dual-running across environments, integrations via EventBridge, REST APIs and IaC with CloudFormation.

    Spark (Scala/Python)DuckDBPandasPolarsPySparkEMREC2LambdaSQSMSKMWAAS3IcebergDelta LakeAthenaRDS/AuroraDynamoDBDocumentDBGlue CatalogDatadogGrafanaCloudWatch
  2. 2025 — 2026

    Data Engineer III · Vert Capital (via Stech Soluções)

    Maintained a Databricks ecosystem with DLT, Unity Catalog and Delta Lake ensuring 90% SLA availability. Optimized batch PySpark pipelines, reducing processing time by 40% through query optimization. Implemented a new end-to-end pipeline using PySpark, Kafka and APIs, collaborating with the Analytics team on data reliability.

    DatabricksDLTUnity CatalogPySparkDelta LakeKafkaPostgreSQLDjangoMetabasePower BI
  3. 2024 — 2025

    Data Engineer II · Stech Soluções

    Developed an end-to-end GCP pipeline focused on data ingestion from multiple sources. Implemented data quality validations with DuckDB in Cloud Functions, and automated processes with Composer (Apache Airflow), reducing manual interventions.

    GCPCloud StorageBigQueryCompute EngineCloud FunctionsComposerPythonDuckDB
  4. 2022 — 2025

    Data Engineer I · SPC Brasil

    Data Lake architect on Hadoop/AWS (20+ TB daily) leading the on-premise to AWS migration. Delivered data quality, lead generation, credit pipelines and ML data serving. Incremental processing of XML/CSV files on-premise and on AWS, with robust orchestration across 100+ daily jobs and REST APIs.

    Spark (Scala/PySpark/Java)HadoopAirflowKafkaNiFiCtrl-MRabbitMQAWS (EC2, EMR, Glue, S3, Athena)HiveImpalaSqoopTerraformMySQLMongoDBOraclePower BIElasticSearch
  5. 2021 — 2022

    Data Engineering Intern · SPC Brasil

    Maintained a Data Lake on Hadoop with data cleaning and partition compaction. Developed Python/Scala scripts for integrity validation and historical reprocessing, and supported Spark infrastructure in staging and production.

    HadoopPythonScalaSpark

Projects

Sparquet

4

Open-source data engineering framework for Apache Spark. Every pipeline is one declarative JSON contract: write it, generate it with any LLM, or design it visually in Sparquet Studio.

TypeScriptapache-sparkdata-engineeringdatabricks

Sparquet Cola

3

Data quality for Spark that runs where your data already is. SODA-style metric checks, SQL rules, schema contracts — plus a valid/invalid split that tells you which rule rejected each row. Pure PySpark, no extra services.

Pythonapache-sparkdata-qualitydata-validation

Pulse

A personal finance dashboard that reads the spreadsheet you already keep on OneDrive. Income and expenses by month, spending by segment and by week, credit card, and a screen just for investments — contributions don't count as spending, yields don't count as income. No database, no new format: it maps the columns you already use.

TypeScriptnextjsreacttailwindcss

Trader Bot

A personal project to place real trades on Binance and help generate some extra income.

Early-stage project — public details still light.

Skills

Languages

Python (PySpark, Pandas, Polars, DuckDB)ScalaJavaSQLCobolRust

ETL / ELT

Spark (Batch/Streaming)GlueEMREC2 (spot)LambdaSqoopDataFusion CometDatabricksSnowflakeCompute EngineCloud FunctionsIncrementalCDC

Messaging

KafkaSQSMSKRabbitMQRedpanda

Orchestration

AirflowMWAAEventBridgeStep FunctionsComposerCtrl-MNiFiDatabricks DLTDatabricks Workflows

Storage & Lakehouse

S3Glue CatalogIcebergHadoopKuduBigQueryCloud StorageDelta LakeUnity CatalogMedallion Architecture

Databases

RDS/AuroraPostgreSQLMySQLDynamoDBMongoDBDocumentDBOracleDB2

Infra

TerraformTerragruntCloudFormation

Analytics & Observability

DatadogGrafanaMetabasePower BIHiveImpalaElasticSearchCloudWatchAthena

Other

CI/CD (Cockpit)Git (GitLab, Bitbucket)Linux (Shell/Bash Script)

Education

  • 2025 – 2026

    Inbix Academy

    MBA — AI for Innovation

  • 2024 – 2025

    FIAP

    MBA — Data Engineering

  • 2020 – 2022

    FATEC São Paulo

    Bachelor's, Systems Analysis and Development

  • 2017 – 2019

    ETEC Jardim Ângela

    Technical, Computer Science — Integrated High School

Contact

I'm open to new opportunities and collaborations — feel free to reach out.