All projects

Multi-source data migration and integration

Adapting processing from Cloudera to Teradata and reconciling multi-source data.

Professional experienceBusiness intelligence data

An assignment from my professional background. No internal or confidential data is disclosed.

Abstract illustration of the main data flow — no internal or confidential data is shown.

Context

Within a Data Factory, I work on data processing, integration and migration between Cloudera and Teradata.

  • Business data from multiple sources
  • PySpark processing on Cloudera
  • Data destined for Teradata

Problem and challenge

Existing processing needs to be adapted to the target environment while translating business rules and bringing multiple sources together.

Objectives

  • Develop PySpark processing on Cloudera.
  • Adapt processing for Teradata with SQL and Shell.
  • Integrate and reconcile multi-source data.

Contribution

I develop and adapt integration processes. I have also built Python processing for data pseudonymisation and reconciliation.

  • Developing PySpark processing and translating business rules.
  • Migrating and adapting processing for Teradata with SQL and Shell.
  • Pseudonymising and reconciling data with Python.

Outcome

Processing adapted for Teradata and data prepared, integrated and reconciled for business needs. The assignment is ongoing.

Technologies

  • Python
  • PySpark
  • SQL
  • Shell
  • Cloudera
  • Teradata
  • Jenkins
  • GitLab

A similar challenge?

Facing a similar challenge? Let’s talk.

Let’s discuss your data, processes and expected outcome to define an engagement that fits your context.

Discuss your project