Sardine is hiring a Senior Data/ML Engineer to own the data and machine learning foundation behind its compliance decisions. The role owns the onboarding pipelines end to end, from data arrival through features to models; it is remote, from the United States or Canada.
Stack
Responsibilities
- ▹Own the data ingestion layer for device telemetry, transaction events, KYC/identity signals and third-party enrichment, designing streaming (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc)
- ▹Build and evolve the feature platform, where the same Chronon feature definitions are computed by Flink for streaming and Spark for batch, with windows from one hour to 300 days, served to the rules engine and models under a sub-second budget
- ▹Establish feature correctness as an engineering discipline: streaming-versus-batch reconciliation, recomputation tests against the warehouse, train/serve parity checks and drift monitoring
- ▹Productionize fraud and identity ML models: training pipelines on Vertex AI and Kubeflow, gradient-boosted and tree-based models (XGBoost, LightGBM, CatBoost, scikit-learn), hyperparameter search, SHAP-based explanations and score normalization
- ▹Build automated retraining, champion/challenger promotion and rollback machinery
- ▹Engineer KYC, AML and identity risk signals: document verification, sanctions/PEP/adverse-media screening, email and phone risk, synthetic identity indicators, bank and account verification, periodic customer due diligence
- ▹Integrate and harden new data sources, including 30+ third-party enrichment providers called in parallel on the request path and the cross-client consortium network, owning failover, timeout budgets, graceful degradation, caching and cost
- ▹Own the warehouse and modeling layer in BigQuery: partitioning, staging-to-mart layers, training datasets and the in-flight migration off dbt onto scheduled SQL and Python pipelines
- ▹Design the entity resolution and graph data linking customers, devices, emails, phones, cards, bank accounts and crypto addresses across clients, including large-scale connected-components work
- ▹Make the platform safe by construction: field-level encryption for sensitive identifiers, regional data residency enforced in pipeline definitions, PII handling and deletion paths, and feature-level gating so a bad signal can be turned off without a deploy
- ▹Set technical direction and raise the team's ceiling: write design docs, run reviews, mentor engineers and data scientists, decide what to build versus buy
Requirements
- ▹8+ years building production data and ML systems, with real ownership of both the pipeline and the model side
- ▹Has shipped models that made consequential automated decisions, not just dashboards
- ▹Deep Python and strong SQL
- ▹Fluency in a distributed processing framework (Spark, Beam or Flink) and streaming semantics: windowing, watermarks, late data, exactly-once versus at-least-once
- ▹Hands-on experience with a modern cloud data stack: GCP strongly preferred (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or AWS equivalents, plus Docker, Kubernetes, Terraform and CI/CD
- ▹Practical ML engineering depth: feature stores and pipelines, training/serving skew, gradient-boosted tree models, class imbalance and rare-event modeling, threshold and cost-sensitive tuning, model monitoring, drift detection and explainability
- ▹Experience with high-volume, low-latency serving where a feature fetch has a few hundred milliseconds and no retry budget
- ▹Domain experience in fraud, risk, payments, lending or identity/KYC, or the demonstrated ability to get fluent in a regulated domain fast
- ▹Understanding of why label latency, feedback loops and adversarial drift make fraud modeling different from ordinary supervised learning
- ▹Comfort with data governance in a regulated environment: PII, encryption, access control, regional data residency, auditability
- ▹Strong written communication, able to explain a modeling tradeoff to a fraud analyst and a pipeline design to a backend engineer
Nice to have
- ▹Experience supporting customer-facing ML: bring-your-own-model integrations, model explainability for adverse action or regulatory review, shadow/challenger scoring frameworks
Soft skills
Strong written communicationExtreme ownership and high growth orientationBias toward action and comfort in ambiguityMentoring and setting technical direction
What we offer
- ▹Work from anywhere in a remote-first culture
- ▹Role based in the United States or Canada
- ▹Hubs in the Bay Area, NYC, Austin, Toronto and São Paulo
- ▹Flexible schedule: performance is valued, not hours worked
About the company
Sardine is the leading agentic risk platform for fighting financial crime, unifying data across risk teams to stop fraud in real time and automate fraud and AML operations. Customers include FIS, GoDaddy, Intuit, Edward Jones, ZoomInfo and Checkout.com.
Ähnliche Stellen

Stelle
Data Engineer Manager
Artefact
AI/ML
+3
💰 Gehalt: keine Angabe
🏢 Vor Ort
🗣️ EN

Stelle
Director, Analytics Engineering (2 Openings)
Novartis
AI/MLBigqueryDatabricksData Science
+5
173 448–322 117 €/Jahr
brutto
🌍 Remote
Position
🗣️ EN
Stelle
Manager, Data Platform
ZoomInfo Technologies LLC
AI/MLBigqueryDatabricks
+8
💰 Gehalt: keine Angabe
🌍 Remote
Anywhere in the World
🗣️ EN

Stelle
Cloud Data Engineer (m/w/d)
Dataciders
DatabricksData Science
+6
💰 Gehalt: keine Angabe
🏢 Vor Ort
Dataciders Standort
🗣️ Deutsch

Stelle
Analytics Engineer
Dandelionhealth
AI/MLData ScienceDbt
+6
124 782–133 695 €/Jahr
brutto
🌍 Remote
🗣️ EN

Stelle
Data Analytics Engineer
Ruby Labs
BigqueryClickhouseDagsterDbt
+5
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN
