hirly

Hubble Network

Data Platform Engineer

San Francisco, CA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Hubble Network first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Stated salary
$153,000 – $250,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
23 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Hubble Network was founded with the intention of delivering on the promise of what Internet-of-Things (IoT) was supposed to be. We're building a global Bluetooth® network dedicated to machine-to-machine connectivity. We differentiate ourselves as the first modem-less and gateway-less, direct-to-satellite network from off-the-shelf Bluetooth® Low Energy chips. Hubble is ideal for applications in logistics, AgTech, and maritime where economies of scale for volume consumer and enterprise asset tracking is a priority. Our goal is to be the first billion-endpoint-connected network in the world.

Hubble is an early-stage, venture-backed startup supported by some of the best investors in the world. In their previous lives, the founding team has been successful in raising $100s of millions in venture funding, developing the Amazon Sidewalk network, launching billions of dollars of space assets, and leading their teams to successful exits, both through acquisition and IPO. We are now looking to bring on talented team members who are the best at what they do to help us make Hubble a reality for the world.

About the Role

This role will be based in San Francisco, CA

This is the founding data hire. You will build Hubble's data platform from the warehouse layer up, make decisions on architecture, determine processes for managing schema and data models, and work with many stakeholders.

We know what we want: a layered lakehouse with Landing, Staging, Warehouse, Mart stages. Each layer has an owner, a purpose, and a quality bar. You’ll work with the Infrastructure team to manage the pipes: Fivetran, Redshift, orchestration, monitoring. You own everything downstream of raw: the transformations, the models, the metric definitions, the quality gates, and the performance bar.

The guiding principle is data democratization. Every person at Hubble should be able to answer their own questions. Success is not how many queries you run for other people, it's how few they need you to run.

You'll be a peer to Platform Engineering and to the team building our internal operations platform. You'll be an embedded consultant to Finance, Product, and Customer Success. You’ll think about how to present a strong story from data.

We expect you to use AI agents as part of how you work, and data work rewards this more than most: schema exploration, model scaffolding, test generation, and documentation are all places where agentic tooling earns its keep if you review the output critically.

Key Responsibilities

Own the transformation layer end-to-end: Defined models across Staging, Warehouse, and Mart. Star-schema design, SCD2, and data contracts enforced by tests at every layer.

Build the data layer our internal tools run on: mart tables designed for specific operational workflows, sitting behind a generated internal metrics API. You own the mart and the API's performance characteristics; Engineering owns the framework and the frontend. Hold the line at p95 under 5 seconds on mart queries and under 500ms on metric endpoints.

Make freshness a commitment, not an aspiration: core data and critical operational metrics streamed in real time with Apache Iceberg, with monitoring that catches breakage before a stakeholder does.

Instrument the funnels: Design event collection across our different product surfaces, unified into a customer 360 spanning authentication, usage, and billing. Acquisition, activation, engagement, retention, revenue.

Data in space: We collect large amounts of telemetry from our satellites which is critical to mission operations - help our Mission Operations group land this telemetry in dashboards and monitoring systems so they can take action on data immediately.

Automate revenue reconciliation: Daily comparison of Stripe Platform billing against Hubble usage, with discrepancies flagged before month-end and drill-down to the device and packet level. Turn a multi-day manual close into a few hours.

Define metrics once: Active devices, packet volume, MRR, contract utilization. Every number has a canonical definition, an owner, and visible lineage from source to dashboard. “Which revenue number is correct?” should stop being a question anyone asks.

Implement classification in the pipeline: Security defines the tiers and the guardrails. You tag datasets by sensitivity, enforce retention, masking, and anonymization, and keep lineage and access logs auditable. Hubble does not store PII, and the pipeline is where that commitment either holds or quietly fails.

Build for self-service: Semantic layers that let business users work with customers and revenue rather than join keys, curated datasets, a data catalog, and documentation good enough that people stop asking you.

Choose boring technology: Fivetran, dbt, Redshift, Metabase are all options - we lean towards buy over build. Bring a strong opinion about which is which, and be able to defend it.

Technical Requirements

Required

7+ years building production data systems, with 2+ years at senior or staff scope owning architecture rather than executing someone else's.

Deep SQL and dimensional modeling: star schemas, slowly changing dimensions, grain discipline. You can explain fact tables and grains.

dbt (or similar) in production: model organization, macros, incremental strategies, tests as contracts, and CI that blocks bad merges.

Cloud data warehouse depth, ideally Redshift (or similar): distribution and sort keys, vacuum and analyze behavior, workload management, and the ability to take a slow query apart and make it fast.

Python or Scala for the specialized pipelines: custom connectors against proprietary APIs, backfills, and reconciliation jobs.

Event and product analytics instrumentation: you've defined a tracking plan, argued about event taxonomy, and built funnel tables that survive contact with a changing product.

Data quality as engineering: anomaly detection, freshness monitoring, and incident response with real escalation paths. You've been paged for a broken pipeline and you've fixed the class of bug, not the instance.

Fluency with AI agents in your own workflow: you already use agentic coding tools to ship real work, and you have a point of view on where they help and where they quietly introduce errors that only show up three models downstream.

Architecture & Scale

Experience with high-volume machine-generated data, telemetry, events, or IoT at billions of rows, where your partitioning and incremental strategy created a cost-efficient and performant data product.

Semantic layer and metric definition work: you've built the abstraction that keeps one number meaning one thing across every tool.

Streaming or near-real-time ingestion (Kinesis, Kafka, Redpanda) with Apache Iceberg or Pinot and a clear view of when it's warranted and when a scheduled batch is the honest answer.

Ownership of a BI tool as a platform: permissions, curated collections, and the discipline that keeps a Metabase instance from becoming an unmaintained graveyard.

Preferred

Interest or experience in IoT, satellite systems, telemetry, or geospatial data.

Billing, usage-based pricing, or revenue reconciliation and recognition workflows.

SOC 2 or similar compliance work: classification frameworks, access reviews, audit evidence.

Open source contributions or public technical writing.

What We Look For

Critical Thinking

You interrogate a metric request before building it, because the question behind it is often different from the question asked. You spot the difference between a data quality problem and a source system problem, and you fix upstream when you can. You make decisions with incomplete information and revise them as you learn.

Partnership

You work as a peer to Platform and Engineering, with clear ownership boundaries and no throwing work over the wall. You sit with Finance and Product long enough to understand what they're actually trying to decide. You

Original posting on Hubble Network's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job