Blog category

AI Features & LLMs

For teams adding AI to a product that has to work in production: which feature is worth building, how LLM, RAG and agent features are put together, what they cost to run, and how to test them so they stay reliable. PadhAI, the AI tutoring platform RAITHub built, is the working example.

19 articles

9 min read· October 2026

Can AI Test Your App? What AI Testing Tools Catch and What They Miss

AI testing tools can now write tests, explore an app in a browser and repair broken selectors. What each kind of tool reliably catches, the five things it misses (your business rules, other users, real money, real devices and real attackers), and a setup that uses AI for speed and a human for judgement.

Read article
13 min read· October 2026

Your AI Wrote the Tests Too: Why They Pass and Still Miss Bugs

AI-written tests usually describe what the code does, not what it should do, so they pass even when the code is wrong. The four patterns behind green-but-useless suites, a five-minute sabotage check, mutation testing with Stryker, and how to stop an agent from editing tests to make them pass.

Read article
10 min read· October 2026

How to QA an App Built with AI Coding Tools

AI coding tools produce code that runs, which is not the same as code that is correct or safe. A practical QA plan for AI-built apps: map the behaviour, test access control first, verify dependencies, check that AI-written tests actually assert something, and put a CI gate in front of every prompt.

Read article
14 min read· October 2026

How to Add AI Features to an Existing SaaS Without Breaking It

A practical plan for adding an LLM feature to a SaaS product that already has paying customers: which feature to start with, where the model call belongs, how to keep cost predictable with routing and caching, how to validate output, handle prompt injection, test and roll out, with lessons from PadhAI.

Read article
12 min read· October 2026

How to Build an AI Tutor Students Can Trust

How to build an AI tutor that helps students learn instead of doing their work: what the research says about unguarded chatbots, how to design a Socratic tutor, how to verify math with code, ground answers with RAG, keep younger learners safe, control cost and evaluate before launch, with lessons from PadhAI.

Read article
13 min read· October 2026

How to Build a RAG Chatbot on Your Own Product Data

A step-by-step build of a retrieval-augmented generation (RAG) chatbot on your own product data: what to index, how to chunk, embeddings and pgvector in Postgres, tenant-scoped retrieval, prompts that cite sources, keeping the index fresh and testing retrieval before launch, with working TypeScript and SQL.

Read article
12 min read· October 2026

Lovable vs Bolt vs Cursor vs Claude Code: Which Survives Production?

An honest comparison of Lovable, Bolt, Cursor and Claude Code for products that must run in production: what each tool actually is, 2026 prices from the vendors’ own pages, who owns and controls the code, a decision table, who each tool is for, and the engineering none of them does for you.

Read article
10 min read· October 2026

RAG vs Fine-Tuning: Which One Your Product Actually Needs

RAG and fine-tuning solve different problems. RAG gives a model facts it can cite at question time; fine-tuning changes how it behaves: format, tone, classification. A comparison for startups covering cost, freshness, data needs, evaluation and when to combine them, with current OpenAI prices and lessons from PadhAI.

Read article
9 min read· October 2026

Is Vibe Coding Bad? When It's Fine and When It Will Cost You

A balanced answer to "is vibe coding bad?": where building by prompt is a sensible trade for a startup, where it turns into expensive rework, what the research and public incidents show, a decision table by situation, and the small set of practices that make it safe to keep going.

Read article
10 min read· October 2026

How to Reduce OpenAI and LLM API Costs by 50–80%

A practical order of operations for cutting an OpenAI or other LLM API bill: measure cost per feature, cap output and reasoning tokens, structure prompts for caching, move waiting work to the Batch API, cache whole responses, trim context and route easy requests to smaller models, with code and a worked example from current list prices.

Read article
11 min read· October 2026

Your AI Chatbot Hallucinates About Your Own Product: How to Fix It

Your support bot quotes a plan that does not exist or a refund policy you never had. How to trace a wrong answer to its cause, whether retrieval, stale content, missing live data or a model filling gaps, and fix it: restrict answers to sources, read prices and limits from the database, verify citations in code, and test with trap questions.

Read article
11 min read· October 2026

RAG vs Agentic RAG: When an Agent Is Worth the Cost

Classic RAG retrieves once and answers once; agentic RAG lets a model plan, search several times, check its own evidence and call tools. How the two differ, a worked cost and latency comparison, the question types that justify an agent, and a bounded hybrid design that sends only hard questions down the agentic path.

Read article
12 min read· October 2026

LLM Model Routing: Send Cheap Queries to Cheap Models

A design guide for LLM model routing: how to choose model tiers, the three router designs (rules, classifier, cascade) and when each fits, a blended-cost formula you can run on your own traffic, how to calibrate a router against an evaluation set, what misroutes cost, and a minimal TypeScript router, with PadhAI's 70/20/10 router as the reference.

Read article
12 min read· October 2026

Preventing Runaway LLM Agents in Production: Limits and Kill Switches

Five guardrails that stop an LLM agent or workflow from looping, burning tokens or running up a bill overnight: a hard step cap, a token budget, a per-tenant spend limit enforced in Postgres, timeouts at call and run level, and a kill switch you can flip without a deploy, with TypeScript and SQL for each and a worked runaway-cost example.

Read article
11 min read· October 2026

How to Add an AI Chatbot to Your Website: 4 Options Compared

The four realistic ways to put an AI chatbot on your website, compared on setup time, cost shape, control, data access and handoff to people: a chatbot widget service, the AI assistant inside a help desk platform, a custom retrieval-augmented (RAG) build, and a hybrid of the two, with worked monthly costs from published prices and a minimal custom endpoint.

Read article
13 min read· October 2026

AI Math Tutors That Don't Give Wrong Answers: Verification Loops

An AI math tutor stops giving wrong answers when code, not the language model, decides what is correct. How to compute truth with a CAS or sandboxed evaluator, check final answers and individual steps, handle "can't verify", and keep answers back in Socratic mode, with a TypeScript checker and lessons from PadhAI.

Read article
14 min read· October 2026

Adaptive Assessment and Knowledge Tracing, Explained for Founders

An adaptive learning engine estimates what each learner knows and picks the next question to match. Here is how rules-based mastery, Bayesian Knowledge Tracing and item response theory compare, what data each needs, and how to handle cold start.

Read article
13 min read· October 2026

What an AI Tutor Really Costs per Student per Month

A worked estimate of AI tutor cost per student per month: tokens per tutoring turn, current list prices from Anthropic, OpenAI and Google, prompt caching, a 70/20/10 model router, RAG and embedding costs, hosting, and a table for light, typical and heavy students, with a TypeScript calculator.

Read article
12 min read· October 2026

AI in a Real Estate CRM: Lead Scoring, Summaries, Follow-Ups

The five CRM jobs where AI pays off for real estate agents, the ones it should not do on its own, how texting consent and fair housing rules shape the design, and what the model calls cost per agent a month on published API prices.

Read article