Maincode
Head of Data
Melbourne
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Lead / management
- Country
- AU
- Work mode
- On-site / unstated
- First seen by hirly
- 30 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Maincode
Maincode makes Matilda, an Australian-made AI stack we own end to end, from tin to token. We own and operate our own infrastructure here in Australia, we train and maintain our own frontier models, and we run multiple surfaces on top of them, including chat, coding and API. Everything we ship is built on capability we control ourselves. We are hiring someone who cares about that. You should be passionate about Australian-made technology, a builder first, and keen to execute at the frontier of data engineering and model training.
The role
A frontier model is only as capable as the data it learns from. We are building a dedicated data organization to source, process, and refine the high value training data that powers Matilda. Our goal is to build the most comprehensive, high quality, and contextually accurate Australian training data catalog in the world.
You will be the foundational leader of this team. You will own the entire lifecycle of our training data catalog. This includes web scale crawling, data licensing, filtering, deduplication, and quality control. We need a technical leader who knows how to recruit top tier data engineers and build a machine capable of turning petabytes of raw information into clean, proprietary training sets.
What you will do
Build and lead a high performing team of data engineers, ML engineers, and data operations specialists.
Design and manage the architecture for web scale data crawling, ingestion, and storage.
Develop robust pipelines for data filtering, cleaning, deduplication, and formatting to ensure the highest quality inputs for model training.
Source and secure proprietary, high value datasets specifically tailored to the Australian context, including navigating licensing and private partnerships.
Work directly with our core training team to iterate on data mixtures, evaluate data quality, and respond to model performance feedback.
Establish the operational standards for data governance, privacy, and security across all our datasets.
What you will bring
Extensive experience leading data engineering or machine learning data teams at scale.
Deep technical knowledge of distributed data processing frameworks and petabyte scale storage systems.
Hands on experience building pipelines for web scraping, data extraction, text processing, and quality filtering.
A proven track record of recruiting, mentoring, and managing elite technical talent.
A strategic understanding of the data landscape, especially regarding sourcing proprietary information and handling complex unstructured data.
Highly regarded
Direct experience building pre-training or post-training datasets for large language models.
A strong network within the Australian technology, academic, or media ecosystem to facilitate unique data partnerships.
Eligibility
Must have the right to work in Australia.
Willingness to undergo background checks as required.
Details
Location: Melbourne, VIC (or Hybrid/On-Site across Australia).
Type: Full time, permanent.
Remuneration: Negotiable based on experience, plus super and equity.
Who thrives here
Leaders who understand that high quality data is the true moat of any frontier model. If you care about Australian-made technology and you would rather spend your week optimizing data mixtures, unblocking your engineering team, and securing high value data sources than sitting in steering committee meetings, you will do well here.
Similar jobs
- Oracle AI Data Platform and Reporting LeadCapgemini · Melbourne, Sydney, AUFirst seen today
- SAP Data Migration Lead/ Consultant - MelbourneCapgemini · Melbourne, AUFirst seen today
- Construction Manager, Data Center Facilities Development (Sydney)Oracle · AustraliaFirst seen today
- Principal Data Engineer (Customer Data & Martech)Cba · 3 LocationsFirst seen today
- Engineering Manager - Database ReliabilityXero · AU: Sydney (45 Clarence St)First seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job