hirly

Rubick.AI

Data Crawling - Consultant

Remote, United States · Bengaluru, India

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Rubick.AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Countries
US, IN
Work mode
Remote-friendly
First seen by hirly
23 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Data Crawling Consultant

Role Overview

You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.

The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements

Part 1: The Fundamentals

  • Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
  • Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
  • Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
  • Monitor crawler performance, identify failures, and resolve data extraction issues.
  • Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
  • Maintain documentation for crawling processes, extraction rules, and data quality standards.

Part 2: AI-Driven Data Crawling Excellence

  • Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
  • Build intelligent crawling workflows using Python and modern web automation frameworks.
  • Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
  • Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
  • Implement automated data validation and monitoring systems to maintain high-quality datasets.
  • Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.

Part 3: Innovation & Improvements

  • Continuously optimize crawling performance for speed, scalability, and reliability.
  • Identify opportunities to automate repetitive extraction and validation workflows.
  • Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
  • Build reusable crawling frameworks and standardized extraction pipelines.
  • Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
  • Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.

To Have

  • 1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
  • Strong proficiency in Python for web scraping and automation.
  • Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
  • Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
  • Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
  • Strong analytical and problem-solving skills.
  • Ability to work with large datasets while maintaining high data quality.
  • Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management is preferred.
  • Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines is a strong advantage.

About Rubick : Company Overview

Rubick OS is a leading platform for brands and marketplaces, enabling seamless management of catalog automation, pricing and competitive intelligence, product listing intelligence, CAST systems, and brand and assortment intelligence.

Our solutions are used across five countries by major brands and marketplaces, including Amazon, Myntra, Reliance Group, Nykaa, Flipkart, Rare Rabbit, Decathlon, Celio, The Bay, Kiabi, Myer, Jumbo, and many more.

Work Model

Work from office – Monday to Saturday

Team Size

250+ Employees

Investors

Invested by Innospark US, Betatron Hong Kong, MJV India, 1Crowd India , and other investors.

Rubick.ai is a place to work if you want to build something foundational rather than incremental—it’s aiming to become the AI operating layer for ecommerce, solving complex, real-world problems across cataloging, pricing, and marketplace operations at scale.

You get high ownership, direct impact, and exposure to various parts of the business, making you an entrepreneur in-house. Rubick operates in the AI segment, which means you will be at the forefront of the technological revolution.

Rubick is profitable and growing and is looking to work with 10,000 brands in the next 2–3 years .

Join us in this exciting journey.

Visit Us

www.rubick.ai

Original posting on Rubick.AI's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job