Micoworks
Senior Data Engineer (Data Platform)
Bangalore
Apply through hirly
hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.
hirly's read of this role
- Role family
- Data & ML
- Seniority
- Senior
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 23 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Senior Data Engineer(Data Platform)
About Mico
Mico's mission is to empower every brand by building lifetime trust through humanlike technology. By 2030, we aim to be Asia's No.1 Brand Empowerment Company. Mico builds the conversational layer Japanese brands use to reach their customers — LINE, SMS/RCS, Voice AI and web — for more than 5,500 companies across finance, insurance, retail and real estate.
Our four core values guide everything we do:
- Wow the Customer
- Invest in Passion
- Beyond Borders
- Be the Change
About the Role
The Data Platform team was newly established in May 2026. Our mission is not a one-off cleanup of the data models in mmc-mono-repo — the codebase behind Mico Engage AI — but building a capability to govern them continuously .
As a Senior Data Engineer, you will own — from investigation through implementation and production rollout — the analysis of entity design and access patterns across TiDB and Aurora, the build-out of our governance toolset (seven items: local reproduction environments, external-expert collaboration, data lineage, test reliability, benchmarks, the classification table, and ADRs), and the concrete work that takes load off TiDB: production deletion under a retention policy, event-attribute transformation, and offloading heavy aggregation to Snowflake with reverse-ETL back into Aurora.
Beyond that lies a succession of core entities such as customer and segment. This is not a greenfield analytics platform role. It is the work of safely reshaping the data model of a large system that is already running, on the basis of evidence .
The Problems We Are Facing
18 applications share one common-lib with 2,500+ imports; the data access layer and business logic are tightly coupled
117 Aurora + 32 TiDB entities, with scattered access patterns and no unified layering definition
Test flakiness is a systemic problem, so coverage numbers cannot be trusted. We believe the root causes lie in parallel execution conflicts, singleton leakage, and Redis state isolation — but identifying and isolating them is itself work still ahead of us
Critical services have grown beyond what one person can hold in their head (CustomerFilterService 61KB, ActionContentSyncService 122KB)
Optimization decisions lack a quantitative basis; performance issues are judged by experience
TiDB has no headroom; the on-call rotation and recovery scripts are what hold it
Where We Are (August 2026)
Six of the seven governance tools have already been through a real case: retention used the benchmark and lineage, and the Snowflake migration used validation cross-checked against production. Lineage layer 1 is past halfway; the raw material for the classification table (the R&R sheet, lineage results, hot/cold analysis) exists and is at the stage of being assembled into one table; Benchmark V1 was used in the retention rehearsal; the monthly meeting with Snowflake is running; and our decision records (Investigation / RCA / Runbook / DesignDoc) are in daily use.
The one remaining item is test reliability , and it is the critical path for everything from Phase 2 (the customer redesign) onwards. Root causes must be listed by October 2026 and fixed by January 2027 — this will be the top priority after you join.
The state we want to be in a year from now (July 2027): TiDB capacity risk confirmed cleared with data; the expiry schedule for legacy data derivable from the SLA; layering as the default call for new requirements; and new members able to get productive on their own from lineage docs and runbooks.
What This Role Is Not
Not "the person who rewrites commonlib"
Not "the person who finishes tuning TiDB in the short term"
Not "the person who moves data to Snowflake and calls it done"
What We Will Not Do
These are explicit team anti-goals.
We will not perform large-scale data migration or schema refactoring before the tools are ready
We will not bypass test reliability and discuss coverage directly
We will not make "performance optimization" decisions without benchmark data
We will not accept "I think this is faster" as an optimization rationale
We will not let an external expert's judgment stop at a single meeting — it must be captured in an ADR
We will not start short-term projects to look busy
We will not over-polish the tools themselves before the toolset's "minimum viable" bar is met
What You Will Do
Analyze data models and access patterns : Build entity-level read/write matrices and lineage ("input entity → business action → output entity") across TiDB and Aurora, and check the results into the repository as mermaid / d2 diagrams with generation scripts.
Layering and storage selection : Using the classification table (transactional/analytical, PII/non-PII, hot/cold, business owner) and lineage as evidence, decide which entities stay in TiDB and which move to Snowflake, and implement the migration. Every decision gets written down.
Codify and execute data retention and deletion : Put the 15-month retention period (a company-level commitment originating in the SLO Project) in writing as a retention SLA and bring it into force; design and rehearse production deletion; execute within the already-fixed window (22:00–24:00, avoiding the delivery calendar and the nightly batch); and turn it into routine operations.
Event-attribute transformation : Move 989 of 1,733 segment conditions (57%) that currently scan raw events onto customer attributes (ever / last timestamp / counter), making those events deletable. An attribute can carry whether , not which , so the 744 conditions that name a specific message or questionnaire keep reading events. The real cost sits not in the pipeline but on the migration side — rewriting 989 segment definitions — and requires alignment with CE / CS and product.
Offload heavy aggregation : Reproduce and validate daily aggregation on the Snowflake side, and build the reverse-ETL pipeline that writes results back into Aurora — preserving the architectural rule that the application keeps reading Aurora as before. The biggest unknown before go-live is the production wiring for writing back into Aurora — IAM, networking, and credentials (production writes are currently manual and approval-gated).
Benchmarking and query optimization : Establish local baselines and QA-environment quantitative baselines for critical query paths, paired with production-side profiles (slow queries, p99, hot-entity access distribution). You will also own slow-query fixes (for example, a duplicate-segment-condition fix measured at roughly 13× improvement) and the SOP for adding indexes on large tables.
Restore test reliability : Identify and fix the root causes of flakiness, then inventory and supplement coverage on critical paths — in that order: fix flakiness before discussing coverage. Root causes must be listed by October 2026 and fixed by January 2027.
Maintain local reproduction environments : Keep and improve the cloned production TiDB environment and the QA event-data generator (both in routine use), and fold the standard build procedure into onboarding.
Incident response and knowledge capture : Take part in the on-call rotation and capture records using our Investigation / RCA / Runbook / DesignDoc templates. Because TiDB has no headroom, when an incident hits, recovery takes priority over the schedule. Convert every incident into a requirement for tool building.
Collaborate with external experts : Join the monthly review with Snowflake, collaborate with PingCAP on a per-design-project basis, and land the important judgments in ADRs.
What We Expect
Take full ownership of the entities and tooling items in your scope — from investigation through architecture, implementation, production rollout, and steady-state operations.
Base optimization decisions on measurement and lineage rather than experience and intuition, and write down the reasoning.
Align directly with the business side (CE / CS) and product, and land c
Browse similar roles
Is this role actually a fit for you?
hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.
Score it against my resume