neuralwithmohit
  • About
  • Blog
  • AI System Design
  • Cloud Comparison
Subscribe
neuralwithmohit

Content

  • Blog
  • AI System Design
  • Case Studies
  • Cloud Comparison

Site

  • About
  • Pricing
  • Privacy

Elsewhere

  • LinkedIn
  • YouTube
  • HuggingFace

© 2026 Mohit Kumar Dubey. Built with Next.js & FastAPI.

neuralwithmohit

On this page

  • Intro
  • What I do
  • Competencies
  • Expertise
  • Career
  • Selected work
  • Products
  • Writing
  • Speaking
  • Contact
Mohit Kumar Dubey

Kuala Lumpur, Malaysia

Mohit Kumar Dubey

Head of AI, ML & Engineering · Entermind Malaysia

I build AI-first products, and the engineering organizations that ship them.

LinkedInSee a case study →
35+

engineers led across AI/ML, backend, mobile & web in a 500+ employee company

3×

delivery output at flat headcount via company-wide LLM automation

45%

lift in AI brand citation scores (GEO platform, GXS · Grab Group)

25%

loan-disbursal conversion uplift (BuddySmart agentic RAG)

30–40%

cut in LLM API cost via model routing & caching (DailyTatva RAG v2)

23%

reduction in lending default rates (credit-risk models)

Leadership

What I do

I lead AI and engineering at the level where a roadmap has to survive both a board conversation and a code review. Over 16+ years I've built and scaled engineering organizations from a handful of people to 35+ across AI/ML, backend, mobile and web, inside a 500+ employee company. I own engineering P&L (licenses, hiring, appraisals) and set the operating model that lets teams ship enterprise AI predictably: technical review gates, architecture sign-off, and delivery prioritization across concurrent engagements.

I carry that from strategy to production. At VerSe Innovation, India's largest short-video and news platform at 350M+ users, I owned the AI/ML roadmap across content and lending and turned it into business outcomes: a 25% loan-conversion uplift, 23% lower default rates, and 30–40% lower inference cost through model routing and cost governance. Today I run AI, ML & Engineering at Entermind, delivering enterprise and government AI across SEA and MENA, where I drove 3× delivery output at flat headcount and act as the primary technical partner to executives from a bank's Chief Credit Officer to the Grab Group.

I still design the systems myself: agentic RAG platforms, multimodal voice digital twins, and an open-source guardrail LLM. I care about the difference between a demo and a system that holds at scale, which comes down to unit economics, AI governance, evaluation before deployment, and the failure modes that only show up in production.

How I operate

Core competencies

Leadership

I design engineering organizations and the technology portfolio behind them. That means scaling teams across AI/ML, backend, mobile and web, developing the people in them, and leading cross-functionally from strategy through delivery.

Business

I own engineering P&L and treat AI product strategy as a commercial lever: supporting enterprise sales, enabling go-to-market, and building the strategic partnerships that turn capability into revenue.

Governance

I set the AI operating model (responsible-AI practice, governance gates, and enterprise adoption paths) and manage executive stakeholders from Chief Credit Officers to group leadership.

Expertise

Technical depth

AI & ML Engineering

  • LLMs, RAG & Agentic RAG
  • Multi-Agent Systems
  • Fine-tuning (LoRA, PEFT)
  • Evaluation & Guardrail Harnesses
  • Prompt Engineering & Structured Outputs
  • Recommendation & Credit-Risk Models

Agent & Serving Stack

  • LangChain · LangGraph · CrewAI
  • vLLM · SGLang · FastAPI
  • LiteLLM gateway
  • PyTorch · TensorFlow
  • Voice: ElevenLabs, Tavus, LiveKit, Pipecat

Cloud & Platforms

  • AWS Bedrock · SageMaker
  • Vertex AI · Azure AI Foundry
  • OpenAI · Gemini · Claude · Cohere
  • Docker · Kubernetes · ECS · EC2 · CI/CD
  • pgvector · Qdrant · Chroma · FAISS · Redis

Experience

Career

Entermind Malaysia

9 mos
Jan 2026 – Present
Kuala Lumpur, Malaysia
  1. Head of AI, ML & Engineering

    Jan 2026 – Present · 9 mos

    Lead an 11-engineer team across backend, AI/ML, frontend and QA, and own the company-wide AI & engineering roadmap for enterprise and government clients across SEA and MENA.

    • Drove a 3× increase in delivery output at flat headcount by identifying and automating recurring workflows with LLM pipelines.
    • Improved pre-sales effectiveness 30% and cut claims and HR validation effort 45% by deploying enterprise RAG and conversational AI.
    • Established the enterprise AI operating model: technical review gates, architecture sign-off, and delivery prioritization across concurrent engagements.
    • Built a reusable GenAI capability layer (RAG ingestion, evaluation harness, guardrail scoring) so engagements start from proven assets, not scratch.

Selected work

Systems I've built

Entermind

Mindara, AI Employee for Enterprise Engagement

An always-on conversational AI colleague that replaces the intake form with adaptive follow-up questioning, surfacing real requirements and handing structured briefs to sales.

15% increase in solutioning lead calls

  • LiteLLM
  • Claude Opus
  • Gemini
  • Langfuse
  • ElevenLabs
  • LangGraph
Entermind

Rukun Ready, Guardrail LLM for Policy Validation

Fine-tuned Qwen2.5-32B to score LLM outputs against national policy principles and emit structured JSON compliance verdicts with safe rewrites. The firm's first open-source AI asset.

3K+ HuggingFace downloads

Products I've led

Shipped at scale

Dailyhunt

News & content · 350M+ users

Led AI/ML and platform architecture for India's largest local-language news app — personalized retrieval, reliability, and inference-cost governance.

30–40% lower LLM cost · 30% lower crash rate · 65% faster fixes

View on Play Store

Josh

Short video · 350M+ users

Led engineering across mobile, backend and ML for the short-video platform — streaming SDKs, recommendations, and reliability at massive scale.

99.8% crash-free · 90%+ launch-to-play · 50% faster playback

Knowledge base

What I write about

The proof of expertise is the content. Four bodies of work sit behind this page. Free and premium material live in the same lists, never partitioned off.

Blog→

Focused deep-dives on one concept, like vector databases, agent memory, or LLM gateway architecture, plus interview scenarios worked end to end.

AI System Design: Guidelines→

The methodology: how to approach designing any AI system. Requirements, evaluation harnesses, latency budgeting, failure modes, and when not to use an LLM.

AI System Design: Case Studies→

Complete architectures for real problems like RAG chatbots, email agents, and multi-agent platforms, each walked through in the same fixed order.

Cloud Component Comparison→

Equivalent services across AWS, Azure, and GCP side by side: the hard limits, pricing shapes, gotchas, and a verdict for each component.

Featured

Reading to start with

Blog · Tech

Understanding Vector Databases

12 min read

GuidelinesPremium

Designing an Evaluation Harness Before You Build

18 min read

Cloud Comparison

Lambda vs Azure Functions vs Cloud Functions

9 min read

Contact

Let's talk

Open to VP / Head of AI & Engineering roles.

Also open to consulting, collaboration, or a talk. The fastest way to reach me is on LinkedIn.

Connect on LinkedIn

Leadership & Governance

  • AI Strategy & Transformation
  • Engineering Org Design
  • Engineering P&L Ownership
  • AI Governance & Responsible AI
  • Executive Stakeholder Management

VerSe Innovation

9 yrs 6 mos· 5 roles
Jun 2016 – Nov 2025
Bengaluru, India
  1. Director of Engineering

    Jun 2025 – Nov 2025 · 6 mos

    De facto Head of Engineering for BuddyLoan; scaled the org from 20+ to 35+ engineers within a 500+ employee company.

    • Improved onboarding conversion 35% and cut drop-off 20% by redesigning acquisition funnels and rebuilding end-to-end telemetry.
    • Deployed AI-driven financial insights and offer personalization (PA/PQ) journeys, extending the platform's monetization surface.
    • Stabilized and scaled the BuddyLoan app and backend platform, aligning mobile, web, backend and AI/ML roadmaps behind one delivery plan.
    • Directed 35+ engineers across AI/ML, backend, mobile and web; instituted operating reviews that improved delivery cadence.
  2. Associate Director of Engineering, AI & Platform

    Nov 2023 – May 2025 · 1 yr 7 mos

    Owned the AI/ML roadmap across content and lending for a 350M+ user platform while leading 20+ engineers.

    • Cut lending default rates 23% and improved model accuracy 15% by deploying credit-risk scoring models for lending partners.
    • Increased search-news engagement 15% and cut manual editorial curation with a multi-agent news summarization system on LangGraph.
    • Owned intent identification, prompt optimization and LLM cost governance across all GenAI products, establishing per-feature unit economics.
    • Sustained 99.8% crash-free sessions, 85%+ notification-to-play and 90%+ launch-to-play across Josh and Dailyhunt.
  3. Software Architect, AI & Platform Architecture

    Apr 2021 – Oct 2023 · 2 yrs 7 mos

    Architecture across the Dailyhunt / Josh 350M+ user platforms.

    • Reduced crash rate 30% and crash-fix turnaround 65% by designing anomaly-detection and crash-analysis pipelines across 350M+ user platforms.
    • Built a GenAI-powered application evaluator for crash prevention and a decisioning service validating CTR/CTA against third-party mappings.
    • Drove cross-platform evaluation (React Native, Kotlin Multiplatform) to inform long-run mobile architecture strategy.
  4. Principal Software Engineer, ML & Recommendations

    Apr 2018 – Mar 2021 · 3 yrs

    Production ML for Dailyhunt / Josh.

    • Improved multilingual personalization across 14+ Indian languages for 350M+ users (audio language detection, video similarity).
    • Engineered modular SDKs for networking, ads and video streaming (custom FMP4/M3U8 with caching).
  5. Lead Software Engineer

    Jun 2016 – Mar 2018 · 1 yr 10 mos

    iOS modernization and cross-platform delivery for Dailyhunt.

    • Achieved 99.8% crash-free sessions and 50% video-playback improvement leading iOS modernization (Obj-C → Swift, MVVM-C).
    • Led React Native delivery and Node.js acquisition services, then transitioned into AI/ML.

Torry Harris Business Solutions

2 yrs 2 mos
Apr 2014 – May 2016
Bengaluru, India
  1. Senior Software Engineer, Enterprise & Government Mobility

    Apr 2014 – May 2016 · 2 yrs 2 mos

    Enterprise and government mobility delivery across the GCC.

    • Delivered 14+ enterprise and government platforms and led an entire business unit across teams, pre-sales, design reviews and CMMI compliance.
    • Built services for Bahrain Government, Kuwait Ministries and Dubai Airports.

GlobalLogic

1 yr 6 mos
Nov 2012 – Apr 2014
Nagpur, India
  1. Senior Software Engineer

    Nov 2012 – Apr 2014 · 1 yr 6 mos

    SmartTV, Xbox and Windows 8 platform applications.

    • Delivered a Verizon movie-rental kiosk platform and a banking-grade OTP encryption SDK (Gemalto Singapore, onsite).

Levitate Mobile Technologies

6 mos
May 2012 – Oct 2012
Gurugram, India
  1. Software Engineer

    May 2012 – Oct 2012 · 6 mos

    Consumer apps for Jubilant Food and Religare Wellness.

Ritusha Consultants

2 yrs
May 2009 – Apr 2011
Kanpur, India
  1. Corporate Trainer

    May 2009 – Apr 2011 · 2 yrs

    Trained 100+ engineers across TCS fresher batches and IIIT-G cohorts.

Full project breakdowns, skills & education on About →
  • Qwen2.5-32B
  • LoRA
  • PEFT
  • vLLM
  • RunPod
  • HuggingFace
Entermind

AI Digital Twin, First Abu Dhabi Bank

A multimodal conversational twin over enterprise knowledge with voice, memory, and agentic workflows, giving stakeholders 24/7 access to the Chief Credit Officer's expertise.

20% lower routine meeting load · 40% better info availability

  • Azure AI Foundry
  • LangGraph
  • Tavus
  • Pipecat
  • pgvector
  • Azure AI Search
Entermind

GEO Platform, GXS (Grab Group)

Citation-aware analytics and optimization across four AI assistants with competitor benchmarking, targeting brand visibility inside AI-driven discovery.

45% lift in AI brand citation scores

  • AWS Bedrock
  • OpenAI
  • Gemini
  • Claude
  • Perplexity
  • LangGraph
VerSe Innovation

BuddySmart, Conversational AI & Agentic RAG

A conversational loan-matchmaking assistant with real-time eligibility checks and personalized offer delivery. It became the platform's primary conversion lever.

25% disbursal conversion uplift · 20% retention gain

  • LangGraph
  • LangChain
  • Cohere
  • Qdrant
  • AWS Bedrock
  • ECS
VerSe Innovation

DailyTatva RAG v1 & v2

Personalized news retrieval at 350M+ user scale. Built the v1 retrieval/inference foundation, then added model routing and caching in v2 while improving retrieval quality.

30–40% cut in LLM API cost

  • LangGraph
  • Gemini
  • Cohere
  • Vertex AI
  • Qdrant
  • ECS
Full project breakdowns on About →
View on Play Store

BuddyLoan

Lending · de facto Head of Engineering

Scaled the engineering org from 20+ to 35+ and owned the AI/ML roadmap — conversational loan matchmaking and credit-risk modeling.

25% disbursal-conversion uplift · 23% lower defaults · 20% retention

View on Play Store

Speaking & video

Neural, with Mohit

I break down AI engineering (RAG, agents, guardrails, and cost control) on my YouTube channel. My latest:

View all on YouTube →
Master DeepAgents: Autonomous Tool Calling, Subagents, Middlewares & Autonomous Workflows

Master DeepAgents: Autonomous Tool Calling, Subagents, Middlewares & Autonomous Workflows

Fixing RAG Retrieval: How to Implement Contextual Chunking & Summarisation - Contextual Retrieval

Fixing RAG Retrieval: How to Implement Contextual Chunking & Summarisation - Contextual Retrieval

LLM Caching with Redis + Qdrant | Cut API Cost & Latency Fast

LLM Caching with Redis + Qdrant | Cut API Cost & Latency Fast