Data Engineer (Danang Onsite)
Đà Nang, Da Nang City, Vietnam
Mô tả công việc
Position Overview:
As a Data Engineer at Madison Technologies, you are the architect of data foundations. You will design, build and maintain the scalable pipelines and warehouse architectures that power our analytics and AI solutions. You will ensure that data flowing from diverse client sources from structured SQL databases to NoSQL backend system such as MongoDB – is seamlessly ingested, transformed, and ready for high-stakes decision-making.
Key Responsibilities:
1. Data Architecture and Pipeline Engineer
• Data Pipeline Development: Redesign ingestion from sales (Pancake) and ad platform APIs (Meta and others), including restatement handling, and cut order and ads latency to 15 minutes or less using incremental MERGE by partition and materialized marts, without doubling BigQuery scan volume.
• Infrastructure and Platform: Build staging, intermediate and mart layers with dbt Core on partitioned, clustered BigQuery tables, and move Airflow from a single VM to a resilient, Terraform-defined setup with an external metadata DB, CI/CD via GitHub Actions, and separate dev, staging and prod environments.
• Data Quality and Governance: Implement dbt tests and Elementary for freshness, uniqueness and reconciliation checks, with alerts for failed jobs and late data routed to Lark through Cloud Monitoring; apply PII policy tags, DLP masking, audit logs and Dataplex catalog and lineage.
• Regulatory Compliance: Implement row access policies for multi-tenant isolation across 9 markets, plus snapshots and DR copies to meet the agreed RTO and RPO targets.
• Cost Optimization: Analyze 90 days of INFORMATION_SCHEMA.JOBS to trace every cost line above 5% of the bill, produce a savings plan, and build a 3-year TCO comparing GCP-only, hybrid (BigQuery plus Iceberg lakehouse) and dedicated/on prem options (ClickHouse, PostgreSQL).
2.Systems Integration and Backend Collaboration
• Cross-System Connectivity: Evaluate and connect additional sources (marketplaces, website/GA4, carriers, COD, bank, Odoo) using Airbyte, dbt or custom Python connectors, normalizing complex data structures into the warehouse.
• Data Activation: Inventory documents across Lark, Google Drive and NAS, run multilingual OCR, and build the semantic index and structured extraction (expiry dates, parties, certificate scope) behind permission-aware search, expiry alerts and future RAG/AI assistant use cases.
Qualifications:
• Experience: 3+ years in data engineering or software engineering, or with a bachelor's degree in a technical field.
• Tech Stack: Strong proficiency in SQL and Python; Go or Java is a plus.
• Orchestration Tools: Hands-on experience with Apache Airflow in production for managing complex task dependencies and scheduling; Terraform and GitHub Actions CI/CD is a plus.
• Ingestion Platforms: Proficiency with automated ingestion tools like Airbyte, Fivetran, or custom API extractors with incremental loads, plus dbt Core and data testing (dbt tests, Elementary or similar).
• Backend Knowledge: Understanding of Data Vault and dimensional modeling and multi-tenant design; exposure to OCR (Document AI or similar), vector search, RAG pipelines or Odoo integration is a plus.
• Cloud Warehousing: Proven experience designing architectures on Google BigQuery (partitioning, clustering, incremental models, cost tuning); Dataplex, policy tags, DLP, Snowflake, Databricks, Iceberg or ClickHouse is a plus.
• Linguistic Proficiency: Good command of English (both written and verbal), ability to present data infrastructure to stakeholders. Vietnamese is a strong plus.
• Location: Based at our Da Nang office.
Why Join Madison Technologies:
• Global Reach: Building infrastructure for diverse global clients across the US, EU and Australia, navigating high-traffic and complex data environment.
• Product-Led Engineering: Work directly with AI specialists and full-stack developers in cross-functional squads, seeing your infrastructure power real-world digital products.