Data Architect + AI

Remote Posted 18h Leaves the board in 7 days
ATS Keyword Match See which key terms from this posting your resume already has. Free.

Data Architect + AI

Overall Stack: Databricks Lakehouse Platform, Apache Spark, Delta Lake, Unity Catalog, and modern cloud data architecture

Must Have

Lakehouse & Medallion Architecture: Expertise in designing end-to-end data architectures (Bronze, Silver, Gold layers) for reliable, production-ready pipelines

Databricks & Spark Internals: Deep understanding of distributed computing. They must know how to troubleshoot and tune large-scale Spark jobs using caching, partitioning, and broadcast joins

Delta Lake: Must understand ACID transactions, schema enforcement, time travel, and optimization operations like Z-ordering

Unity Catalog & Data Governance: Proven ability to design unified governance models for data and AI assets, including role-based access control (RBAC), row/column-level security, and data lineage

Cloud Infrastructure (AWS, Azure, or GCP): Strong grasp of the native cloud ecosystem they work in (e.g., ADLS/Entra for Azure, S3/IAM for AWS), including Virtual Network (VNet) setups and IAM roles

Coding Proficiency: Advanced SQL skills and fluency in Python or Scala

Cost Optimization & Performance Tuning: Ability to monitor DBUs (Databricks Units), right-size serverless and multi-node clusters, and implement best practices for avoiding cloud bill shock

Nice to Have

Databricks Certifications: Candidates holding valid Databricks Certified Data Architect or Databricks Certified Data Engineer Professional badges generally have a proven, up-to-date baseline of the platform's features

Generative AI & MLflow Integration: Experience building, deploying, and monitoring GenAI applications and ML models using Databricks Model Serving, Vector Search, and the Mosaic AI suite

CI/CD & DevOps Practices: Experience automating Databricks workflows using Git (Databricks Repos) and orchestration tools like dbt, Azure Data Factory, or Apache Airflow

Streaming Data: Familiarity with Databricks Structured Streaming and Auto Loader for real-time data ingestion and processing

Data Warehousing & BI: Understanding of Databricks SQL, Serverless Warehouses, and integration with downstream BI tools like Power BI

Remote

Adavenced english

Originally posted on Himalayas

Share your thoughts