All 72 articles, sorted alphabetically
Advanced RAG, in depth: HyDE, multi-query, step-back, decomposition and GraphRAG as a routed query-transformation layer
A practical guide to advanced RAG as a prompt layer: why verbatim queries miss, the prompts for HyDE, multi-query, step-back and decomposition, what G…
Read article →Agentic Prompting, in depth: task briefs, autonomy boundaries, progress state and verification for long-running agents
How to write the prompts that drive autonomous, multi-step agents: what changes from single-turn prompting, the anatomy of a task brief, checkable don…
Read article →Chain-of-Thought (CoT) Prompting, in depth: putting reasoning into production with output contracts, token budgets, routing, hidden chains and distillation
Chain-of-thought as a production component rather than a prompt trick: the reasoning-before-answer output contract, token and latency budget arithmeti…
Read article →Constitutional AI as a prompting pattern, in depth: writing principles, critique and revision loops, and AI feedback you can run yourself
How to apply the ideas of Constitutional AI in your own prompt pipelines: what the original method did during training, how to write a constitution of…
Read article →Context Window Management, in depth: token budgets, output reserves, an overflow ladder for long conversations and agents, and a cache-stable prefix
How to manage a model's context window at run time: what counts against it, counting tokens with the model's own tokenizer, reservin…
Read article →Cost Optimization, in depth: a per-request cost model, cost per successful task, and the prompt levers in the order that pays
How to make LLM prompts cheaper without losing quality: compute cost from usage fields, measure cost per successful task, then apply caching, input hy…
Read article →Emergent Abilities, in depth: what scale-dependent behaviour means for prompt engineers, how to measure it on your models, and how to port prompts to smaller ones
A practitioner's guide to emergent abilities in language models: the original definition and the meas…
Read article →Dynamic Few-Shot, in depth: algorithms for choosing in-context examples per request, from kNN and MMR to label-stratified and learned selection
How to select few-shot examples per request: what selection can and cannot change, building the example pool, what to embed, top-k nearest neighbours,…
Read article →Few-Shot Prompting, in depth: what demonstrations actually teach, the biases they add, calibration, and how to measure each example
Few-shot prompting from first principles: how in-context demonstrations condition a model, what the research says they really carry (format, label spa…
Read article →Function Calling, in depth: writing tool definitions models can use, the call loop, tool_choice, results, safety and evaluation
The prompt-engineering side of function calling and tool use: how the protocol works, anatomy of a tool definition in the Claude and OpenAI APIs, writ…
Read article →Grounding + Citations, in depth: source packing, the citation contract, verbatim-quote validation and abstention
How to make an LLM answer only from supplied sources and prove it: what grounding and citations do and do not guarantee, packing sources with stable I…
Read article →Instruction Hierarchy, in depth: privilege levels in LLM applications, writing system prompts that hold, treating tool output as data and testing conflicts
What the instruction hierarchy is and where it comes from, how system, developer, user and tool content rank, how to write system prompts that declare…
Read article →Long Context Prompting, in depth: document layout, query placement, quote-first answers and measuring where your model stops reading
How to prompt a model with tens or hundreds of thousands of tokens of input: why a bigger window is not a bigger memory, where to put documents and qu…
Read article →Model-Specific Prompt Tricks, in depth: what really differs between Claude, OpenAI reasoning models and Gemini, and how to keep prompts portable
Model-specific prompting from first principles: the documented differences between Anthropic, OpenAI reasoning-model and Gemini guidance as of October…
Read article →Multi-Agent Orchestration, in depth: briefs, handoffs and synthesis prompts that hold up
How to orchestrate several LLM agents: when splitting helps, topologies, orchestrator and worker briefs, handoff schemas, synthesis and verification p…
Read article →Negative Prompting, in depth: why do-not instructions misfire, how to rewrite them, and how to enforce bans outside the prompt
A practical guide to negative instructions in LLM prompts: why negation is fragile for language models, positive reframing with worked rewrites, when …
Read article →Agent Prompt Architecture in Depth
Prompting an agent rather than a single-turn model: the system prompt as specification, tool descriptions and parameter naming, choosing between tools…
Read article →Chain-of-Thought Prompting, in Depth: Why Reasoning Tokens Help, How to Elicit and Parse Them, When They Hurt, and How to Measure the Gain
A practical guide to chain-of-thought prompting: why generating intermediate steps improves multi-step accuracy, zero-shot and few-shot patterns, answ…
Read article →Chain-of-Verification Architecture, in depth: planning, factored verification and revision as a production pipeline
How Chain-of-Verification (CoVe) reduces hallucination: the four stages from the 2023 Meta AI paper, why joint, two-step, factored and factor+revise v…
Read article →Context packing architecture
How to assemble the highest-value tokens into a fixed LLM context window: the anatomy of a context pack, salience scoring and reranking, token budgeti…
Read article →Dynamic few-shot prompting
Deep-dive on dynamic few-shot prompting, where the demonstrations shown to a model are selected per request from an exemplar store rather than frozen …
Read article →Prompt Evaluation Architecture in Depth
Evaluating prompts and LLM systems: representative eval sets, the three scorer families, LLM-as-judge bias and validation, pairwise comparison, sampli…
Read article →Few-shot prompting architecture, in depth: designing, formatting, ordering, caching and evaluating a fixed example set
Few-shot prompting treated as a system: why in-context demonstrations work, what an example set must cover, inline tags versus multi-turn formatting, …
Read article →LLM hallucination guardrails architecture
Deep-dive on layered hallucination guardrails for LLM systems: why fabrication is structural and trust asymmetric, grounding via retrieval, the cite-o…
Read article →Least-to-most prompting architecture
Deep-dive on least-to-most prompting: decomposing a hard problem into an ordered easiest-to-hardest subproblem queue, passing each committed subanswer…
Read article →Meta-prompting -- using an LLM to write and improve prompts
Deep-dive on meta-prompting: using an LLM to generate and refine prompts, the optimization loop (generate/evaluate/refine), the meta-prompt, automatic…
Read article →Multimodal prompting architecture
Deep-dive on prompting with images: how pixels become tokens via resize and tiling, low vs high detail cost, image placement and interleaving, the cro…
Read article →Output Formatting, in depth: choosing and specifying the shape of model output for people, parsers and renderers
How to decide what shape an LLM's output …
Read article →Output Parsing Architecture in Depth: Turning Model Text into Trusted, Typed Data
How to build the parsing layer between an LLM and your code: classifying how a response ended, extracting payloads from free text, a bounded tolerant-…
Read article →Prompt caching architecture
Deep-dive on LLM prompt caching: storing the deterministic KV attention state of a stable prompt prefix so later requests prefill only the new tail, t…
Read article →Prompt compression architecture
Deep-dive on compressing LLM prompts at scale: segment inventory and per-class token budgets, relevance scoring, the extractive-abstractive-pruning ca…
Read article →Prompt-injection defense architecture
Deep-dive on defending LLM applications against prompt injection: why no prompt can stop it, how to separate trusted instructions from untrusted conte…
Read article →Prompt Pipeline Architecture in Depth
A 2500-word walkthrough of a production prompt pipeline: templates, variables, compiler, structured outputs, model router, validation, registry, eval,…
Read article →Prompt registry architecture
Deep-dive on prompt registries: immutable prompt versions with variable schemas, eval-gated promotion through environment labels, cached runtime resol…
Read article →Prompt Templates, in depth: typed contracts, strict rendering, token budgets, cache-friendly ordering and golden-render tests
How to design prompt templates as software: typed input contracts, strict and sandboxed rendering, the double-render injection bug, delimiting untrust…
Read article →ReAct prompting
Deep-dive on ReAct prompting: the thought-action-observation loop, tool grounding for reduced hallucination, prompt format vs native tool calling, err…
Read article →Reflexion architecture
Deep-dive on the Reflexion pattern: how an LLM agent improves within a single session through verbal reinforcement — writing natural-language self-cri…
Read article →Role Prompting, in depth: what a role actually changes, what the evidence says about personas, writing role blocks as specifications, and testing them with an A/B harness
A practical guide to role prompting: how a role conditions a model, what research on personas in system prompts found about accuracy, the four things …
Read article →Prompt routing architecture
Deep-dive on prompt routing: intent classifier, policy engine, LLM registry, selection, fallback ladder, quality gate, and A/B.
Read article →Self-consistency -- sample many reasoning paths, vote
Deep-dive on self-consistency: sampling multiple chain-of-thought reasoning paths and taking the majority-vote answer, marginalizing over the reasonin…
Read article →Semantic routing architecture
Deep-dive on semantic routing for LLM apps: embedding queries, labeled route centroids, cosine similarity and confidence thresholds, LLM fallback for …
Read article →Skeleton-of-Thought prompting architecture
Deep-dive on Skeleton-of-Thought: a cheap skeleton call that lists an answer&a…
Read article →Step-back prompting architecture
Deep-dive on step-back prompting: the two-stage abstraction-then-reasoning pipeline, using the derived principle as a retrieval key, self-verification…
Read article →Structured Output Architecture for LLMs in Depth
A 2500-word walkthrough of structured output: schema, provider features (function calling, JSON mode), constrained decoding, validation, retry, stream…
Read article →Tree of Thoughts architecture
Deep-dive on Tree of Thoughts prompting: thought decomposition, candidate generation, value vs vote state evaluation, BFS/DFS search with beam width a…
Read article →Verifier Architecture, in depth: separating generation from checking, verifier types, calibrated thresholds, best-of-N, reject-and-retry and fallbacks
How to design an LLM verifier layer: why checking is easier than generating, programmatic checks, model graders and outcome versus process reward mode…
Read article →Zero-Shot Prompting, in depth: Writing Instruction-Only Prompts as Specifications, Evaluating Them, and Knowing When to Escalate
Production zero-shot prompting: the five parts of an instruction-only prompt, label definitions and output contracts, zero-shot chain of thought, vali…
Read article →Prompt A/B Testing in Production, in depth: assignment per user, exposure logging, judge metrics, sample size, the delta method and prompt-specific confounders
How to run online A/B tests of LLM prompts: where they sit after offline evals, randomizing by user or conversation, logging prompt hashes and exposur…
Read article →Analogical Prompting, in depth: letting the model write its own exemplars, when it helps and how it fails
How analogical prompting works: the model recalls relevant and distinct problems with solutions before solving the target, as proposed in Large Langua…
Read article →Anatomy of a Prompt, in depth: what the model actually receives, the components and their jobs, role authority, ordering and ablation
A first-principles breakdown of an LLM prompt: how chat templates flatten messages into tokens, the components a production prompt is built from and w…
Read article →Prompt Chaining, in depth: when to split a task, step contracts and validation gates, error-compounding arithmetic, and a worked support-ticket chain
How to design, build and debug a prompt chain: when a chain beats a single prompt, the common topologies, step contracts and deterministic gates betwe…
Read article →Prompt Compression with LLMLingua, in depth: perplexity-based token pruning, question-aware LongLLMLingua, the LLMLingua-2 classifier and when compression pays
How the LLMLingua family compresses prompts: why token information is uneven, Selective Context, LLMLingua's budget controller and iterative toke…
Read article →Debate Prompting, in depth: stance prompts, rebuttal rounds, judge rubrics and the baseline debate has to beat
A practical guide to the prompts behind multi-agent debate: independent and stance-assigned openings, the round-N rebuttal prompt with anti-conformity…
Read article →Debugging Broken Prompts, in depth: capture, reproduce, minimise, classify, fix one thing and lock it in
A systematic workflow for debugging LLM prompts that misbehave: capture the exact rendered request, measure the failure rate across runs, minimise the…
Read article →Prompt Delimiters
Separate instructions from data using delimiters: XML tags, markdown headers, or triple quotes. Reduces injection risk and improves model comprehensio…
Read article →Prompting for Code Generation, in depth: repository context, tests as the specification, diff contracts, the execute-and-repair loop and pass@k evaluation
How to prompt language models to write code that survives review: what context to assemble from the repository, using tests as the specification, choo…
Read article →Domain-Specific Prompts, in depth: text-to-SQL prompts built from your schema, your definitions and your verified queries
How to prompt a model to write correct SQL against your own database: representing the schema, schema linking, a business glossary, retrieved verified…
Read article →Prompt Engineering Team Structure, in depth: the work that exists, three operating models, roles, ownership in the repo and model migrations
How to organise people around prompts and LLM features: an inventory of the real work, centralised, embedded and hub-and-spoke models, a roles and res…
Read article →Extraction Prompts, in depth: evidence quotes, absent versus unclear, normalising in code, long documents and per-field evaluation
How to write and run prompts that extract facts from documents reliably: field definitions that remove ambiguity, exact-quote evidence verified agains…
Read article →Multilingual Prompting, in depth: pinning the output language, glossaries, locale formatting in code and per-language evaluation
How to build prompts that work across languages: choose the instruction language, pin the output language and register, protect product terms with a g…
Read article →Prompt Design for Streaming UX, in depth: answer-first prompts, stream-safe formats, markers, partial JSON, buffered moderation and cancellation
How to design prompts and clients for streamed LLM output: why order matters once users read as tokens arrive, time-to-useful-token, answer-first prom…
Read article →Temperature, Top-p, Top-k, in depth: what each knob does to the next-token distribution, why order matters, what providers allow, and how to tune settings per task
A practitioner's guide to sampling parameters: softmax and temperature worked through with real numbers, top-k, top-p and min-p trunc…
Read article →Prompt Template Libraries
Reusable prompt templates, variables, composition, and library trade-offs. LangChain PromptTemplate, Guidance, Jinja2 for structured prompts that scal…
Read article →Prompt Versioning, in depth: what a version contains, content-addressed releases, model pinning, eval-gated changes and rollback
How to version prompts properly: the full bundle a version must capture, content hashes as identity with aliases for release, change classes, a git an…
Read article →Reflexion, in depth: building the loop, designing evaluators and reflections, and reading what the paper's numbers really show
A practical guide to implementing Reflexion, the pattern where an LLM agent retries a task with written lessons from its own failures: what the origin…
Read article →RAG, in depth: chunking, hybrid retrieval, reranking, packing the prompt, evaluating each stage and the failure modes
Retrieval-augmented generation from first principles: why it works, structure-aware chunking with context headers, hybrid dense and lexical retrieval …
Read article →Role Prompting, in depth: persona specs that survive long conversations, drift, simulated users and roleplay attacks
How to design, build and test long-lived personas for LLM products: writing a persona spec, separating persona from policy, why personas drift over ma…
Read article →Self-Consistency, in depth: answer extraction, voting, early stopping and agreement as a confidence signal
Engineer self-consistency for production: why answer extraction and normalisation decide whether voting works, a provider-agnostic parallel implementa…
Read article →Self-Refine, in depth: generate, critique and rewrite with one model, what the evidence shows, prompt design, a working loop and when not to use it
A practical guide to Self-Refine prompting: the generate, feedback and refine loop from Madaan et al. (2023), what the paper measured and where it fai…
Read article →Structured Output, in depth: one JSON Schema contract, compiled for strict modes, validated twice and versioned
Treat LLM JSON output as a versioned contract: one schema source, compiled to the strict subsets that Claude and OpenAI enforce, stop-reason gating, f…
Read article →Task Decomposition, in depth: where to cut a task, plans as validated graphs, and when not to split
How to decompose tasks for language models: seam tests for where to cut, static versus model-generated plans, a JSON plan validated as a DAG, a parall…
Read article →Zero-Shot Prompting, in depth: why instruction-tuned models follow a bare instruction, scoring labels with log-probabilities, calibrating label bias and measuring prompt sensitivity
Zero-shot prompting from the measurement side: why instruction tuning makes instruction-only prompts work, generate-and-parse versus scoring each labe…
Read article →