PointFive
Head of LLM research
Tel Aviv
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- IL
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About PointFive
- PointFive finds the cost buried inside enterprise infrastructure, from the cloud to coding agents. We pioneered DeepWaste across cloud and data platforms, and we are now doing the same for AI token spend with TokenShift.
- Fanatics, H&M, Hertz, Nubank, Citizens Bank and other Fortune 500 companies use PointFive to see and cut what their infrastructure and AI actually cost.
We have raised $96 million, most recently a $60 million Series B led by Accel . Engineering sits in Tel Aviv and our go-to-market team is focused on North America. PointFive Labs, our research arm, publishes original work on how infrastructure and AI actually get paid for.
CRN named us one of the 10 hottest cloud computing startups of 2026 , Forbes Israel put us on its Next Billion Dollar Startups list , and Redpoint selected us for its InfraRed 100 , its ranking of the private companies shaping the future of cloud infrastructure.
About the Role
PointFive is building infrastructure for the next generation of AI-powered engineering.
As AI agents become embedded into developer workflows, the underlying model layer is changing rapidly. Organizations are no longer choosing between a handful of hosted APIs. They are increasingly operating across frontier models, open-weight models, locally deployed models, specialized models, and dynamically routed combinations of them.
We’re looking for an AI researcher to take an integral part in PointFive’s research into how these models behave, how they should be evaluated, and how they can be deployed and used more efficiently across real engineering workloads.
This is a deeply technical research role with direct product impact.
You’ll study frontier and open-source models, develop evaluation frameworks, investigate inference and model optimization techniques, explore local deployment strategies, and help answer questions such as:
- Which model should handle a particular task?
- How much capability do we lose when we quantize it?
- Can a smaller local model replace a frontier API for a specific workload?
- How should models be routed across latency, cost, privacy, and quality constraints?
- How do we measure whether one model is actually better than another for agentic software engineering?
Your work will directly shape PointFive’s model strategy and the intelligence behind our AI infrastructure products.
What You'll Do
Define and lead PointFive’s LLM research agenda across model evaluation, inference optimization, local models, routing, and model adaptation.
Continuously evaluate frontier and open-weight models across real-world software engineering and agentic workloads.
Design rigorous evaluation frameworks for model quality, reasoning, tool use, code generation, command execution, summarization, and agent behavior.
Build benchmarks that reflect actual developer workflows rather than generic academic tasks.
Research model efficiency techniques including quantization, distillation, speculative decoding, prefix caching, KV-cache optimization, batching, and context management.
Investigate when smaller or locally deployed models can replace expensive frontier models without materially degrading task quality.
Evaluate model architectures, parameter sizes, quantization formats, and runtime configurations across heterogeneous hardware.
Research and benchmark local inference stacks including llama.cpp, MLX, vLLM, SGLang, ONNX Runtime, Ollama, TensorRT-LLM, and emerging runtimes.
Study inference performance across Apple Silicon, NVIDIA GPUs, AMD GPUs, CPUs, Windows workstations, Linux machines, and other endpoint configurations.
Develop model-routing strategies that optimize across quality, latency, cost, privacy, context size, and hardware availability.
Explore intelligent cascades where smaller models handle common tasks and more capable models are invoked only when necessary.
Research model specialization through fine-tuning, LoRA, adapters, distillation, prompt optimization, and other model adaptation techniques.
Investigate model behavior under constrained environments, including offline execution, limited memory, limited compute, and local-only inference.
Evaluate agent-specific model behavior, including planning, tool selection, shell interaction, code editing, error recovery, and long-running task execution.
Analyze failure modes such as hallucination, tool misuse, context degradation, reasoning collapse, excessive token consumption, and unstable agent loops.
Design experiments that quantify the tradeoffs between model capability, inference cost, token consumption, latency, and resource utilization.
Build internal research infrastructure for reproducible model experiments, benchmarking, dataset management, and evaluation.
Track frontier model releases and emerging research, rapidly determining which developments are meaningful for PointFive’s products.
Collaborate closely with engineering and product teams to translate research results into production capabilities.
Build and lead a small, exceptional LLM research team over time.
What We're Looking For
Must-have
Deep understanding of modern large language models and transformer-based architectures.
Strong hands-on experience evaluating and experimenting with both frontier and open-weight models.
Strong understanding of inference behavior, including prefill, decoding, KV caches, context windows, batching, memory usage, and token generation performance.
Experience with model optimization techniques such as quantization, distillation, LoRA, fine-tuning, or model compression.
Strong experimental mindset — able to formulate hypotheses, design controlled experiments, and draw meaningful conclusions from noisy results.
Experience building evaluation frameworks for LLM quality and behavior.
Strong Python proficiency and familiarity with the modern ML ecosystem.
Ability to read, understand, and reproduce ideas from current ML research papers.
Comfortable working with ambiguous research problems where there may not yet be an established best practice.
Strong ability to bridge research and production — understanding not only whether something works, but whether it is practical to deploy.
Nice to have
Experience with open-weight models such as Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, or similar model families.
Experience with frontier model APIs including OpenAI, Anthropic, Google, and other leading providers.
Experience with inference frameworks such as vLLM, SGLang, llama.cpp, MLX, TensorRT-LLM, Ollama, or ONNX Runtime.
Deep understanding of quantization techniques including FP8, INT8, INT4, AWQ, GPTQ, GGUF, and related approaches.
Experience with GPU performance optimization, CUDA, Metal, ROCm, or DirectML.
Familiarity with distributed inference and multi-GPU serving.
Experience with model routing, mixture-of-model systems, cascades, or adaptive inference.
Experience with reinforcement learning, preference optimization, DPO, GRPO, or related post-training techniques.
Familiarity with agentic systems, coding agents, tool-use models, and computer-use models.
Experience building or evaluating coding benchmarks and software-engineering agents.
Experience with synthetic data generation, dataset curation, and automatic evaluation.
Research publications or meaningful contributions to open-source ML projects.
Experience leading a small applied research or ML research team.
Research Areas
Some of the problems we expect this team to work on include:
- Model Routing
- Choosing the optimal model dynamically based on task difficulty, latency requirements, cost, privacy constraints, and hardware availability.
- Local vs. Cloud Inference
- Determining which workloads can reliably move from cloud models to models running directly on developer endpoints.
- Model Compression
- Understanding how far models can be quantized, distilled, or otherwise optimized before meaningful
Similar jobs
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job