Mani Pal

Engineer-researcher

Mani Pal

LLM systems, CUDA kernels, inference optimization, compression, interpretability, and distributed AI infrastructure.

Available · 2 sprint slots this monthReplies within 12 hours, IST (UTC+5:30)

I cut LLM inference cost and latency. Fixed scope. Benchmarked. Guaranteed.

If you serve models in production and the GPU bill or p95 latency hurts, I find the bottleneck and move the number — then hand everything to your team. Every engagement is pinned to a measurable target up front. Miss the target, and you don't pay the second half.

fig. 02 — the work, in one picture

live

Same sweep your serving stack runs millions of times an hour. I make each pass cheaper.

Evidence, not adjectives

Every number links to the full write-up — methodology, profiling traces, and code.

Engagements

Three ways to work together

Inference Audit Sprint

2 days

$1,500 fixed

50% to book the slot · 50% on delivery

I profile your serving stack end to end and hand you a prioritized optimization plan with projected cost and latency savings.

  • Profiling of your current serving path (vLLM / SGLang / TensorRT-LLM / llama.cpp / custom)
  • Bottleneck report: batching, KV-cache, kernels, quantization, routing
  • Prioritized optimization plan with projected savings per item
  • 30-minute walkthrough call with your team

If the audit doesn't identify at least 20% of provable cost or latency improvement, the second half of the fee is refunded.

Speedup Sprint

3–5 days

$2,500–3,500 fixed

50% to book the slot · 50% on delivery

I implement the highest-impact optimization from the audit — or one you already know you need — and prove it with before/after benchmarks.

  • Implementation: batching strategy, quantization, speculative decoding, KV-cache tuning, or kernel-level work
  • Before/after benchmark report on your real traffic shape
  • Production-ready code merged into your repo with tests
  • Handoff doc so your team owns it after I leave

Scope is pinned to an agreed benchmark target up front. Miss the target, and the second half of the fee is refunded.

Fractional Inference Engineer

Monthly · limited seats

From $5,000/month

Billed monthly in advance

Ongoing ownership of your inference layer — cost, latency, reliability — without a full-time hire.

  • Continuous profiling and optimization of your serving stack
  • Architecture reviews for new model launches and traffic growth
  • On-call for inference incidents during agreed hours
  • Monthly cost/latency report to leadership

Month-to-month. Cancel anytime with 2 weeks' notice.

Process

Intro call to handoff

01

Intro call

15 minutes. You describe the stack and the pain; I tell you honestly whether I can move the number.

02

Scope + deposit

Fixed scope and benchmark target in writing. 50% deposit locks the slot; I typically start within a week.

03

Sprint

Heads-down work with a short written update every day. No meetings unless something needs a decision.

04

Handoff

Benchmark report, merged code, handoff doc. Your team owns everything; I stay reachable for follow-ups.

Good fit

  • You serve LLMs in production and the GPU bill or p95 latency hurts
  • You run vLLM, SGLang, TensorRT-LLM, llama.cpp, or a custom serving path
  • You want a fixed-scope engagement with a measurable target, not open-ended hours
  • Your team will own the result — I optimize and hand off

Not a fit

  • No production traffic yet — optimization before load is premature
  • You need a full application built end to end
  • You want indefinite hourly staff augmentation
  • The bottleneck is the model's quality, not its serving cost

FAQ

How do you access our stack?

Read-only repo access plus profiling traces is usually enough for the audit. For implementation sprints, a scoped branch and CI access. I'll sign your NDA before seeing anything.

What about timezones?

I work from Delhi (IST, UTC+5:30) — full overlap with European hours, and US East mornings / US West evenings are covered. Sprints are async-first with a written daily update.

How do payments work?

Wise, wire transfer, or Stripe invoice — whatever your finance team prefers. 50% books the slot, 50% on delivery against the agreed benchmark.

How soon can you start?

Typically within a week of the deposit. If the slot pill above says slots are open, the queue is short.

Available · 2 sprint slots this month

Fifteen minutes to find out if I can move your number.

Bring your serving setup and the metric that hurts. I'll tell you honestly whether there's 20%+ on the table — and if there isn't, I'll say so on the call.