Prompt Engineering

Prompt Engineering

Deep technical articles on this topic.

83Articles
83Topics covered
Articles in this category

All 72 articles, sorted alphabetically

Advertisement
ARTICLE · 01

Advanced RAG, in depth: HyDE, multi-query, step-back, decomposition and GraphRAG as a routed query-transformation layer

A practical guide to advanced RAG as a prompt layer: why verbatim queries miss, the prompts for HyDE, multi-query, step-back and decomposition, what G…

Read article →
ARTICLE · 02

Agentic Prompting, in depth: task briefs, autonomy boundaries, progress state and verification for long-running agents

How to write the prompts that drive autonomous, multi-step agents: what changes from single-turn prompting, the anatomy of a task brief, checkable don…

Read article →
ARTICLE · 03

Chain-of-Thought (CoT) Prompting, in depth: putting reasoning into production with output contracts, token budgets, routing, hidden chains and distillation

Chain-of-thought as a production component rather than a prompt trick: the reasoning-before-answer output contract, token and latency budget arithmeti…

Read article →
ARTICLE · 04

Constitutional AI as a prompting pattern, in depth: writing principles, critique and revision loops, and AI feedback you can run yourself

How to apply the ideas of Constitutional AI in your own prompt pipelines: what the original method did during training, how to write a constitution of…

Read article →
ARTICLE · 05

Context Window Management, in depth: token budgets, output reserves, an overflow ladder for long conversations and agents, and a cache-stable prefix

How to manage a model's context window at run time: what counts against it, counting tokens with the model's own tokenizer, reservin…

Read article →
ARTICLE · 06

Cost Optimization, in depth: a per-request cost model, cost per successful task, and the prompt levers in the order that pays

How to make LLM prompts cheaper without losing quality: compute cost from usage fields, measure cost per successful task, then apply caching, input hy…

Read article →
ARTICLE · 07

Emergent Abilities, in depth: what scale-dependent behaviour means for prompt engineers, how to measure it on your models, and how to port prompts to smaller ones

A practitioner's guide to emergent abilities in language models: the original definition and the meas…

Read article →
ARTICLE · 08

Dynamic Few-Shot, in depth: algorithms for choosing in-context examples per request, from kNN and MMR to label-stratified and learned selection

How to select few-shot examples per request: what selection can and cannot change, building the example pool, what to embed, top-k nearest neighbours,…

Read article →
ARTICLE · 09

Few-Shot Prompting, in depth: what demonstrations actually teach, the biases they add, calibration, and how to measure each example

Few-shot prompting from first principles: how in-context demonstrations condition a model, what the research says they really carry (format, label spa…

Read article →
ARTICLE · 10

Function Calling, in depth: writing tool definitions models can use, the call loop, tool_choice, results, safety and evaluation

The prompt-engineering side of function calling and tool use: how the protocol works, anatomy of a tool definition in the Claude and OpenAI APIs, writ…

Read article →
ARTICLE · 11

Grounding + Citations, in depth: source packing, the citation contract, verbatim-quote validation and abstention

How to make an LLM answer only from supplied sources and prove it: what grounding and citations do and do not guarantee, packing sources with stable I…

Read article →
ARTICLE · 12

Instruction Hierarchy, in depth: privilege levels in LLM applications, writing system prompts that hold, treating tool output as data and testing conflicts

What the instruction hierarchy is and where it comes from, how system, developer, user and tool content rank, how to write system prompts that declare…

Read article →
ARTICLE · 13

Long Context Prompting, in depth: document layout, query placement, quote-first answers and measuring where your model stops reading

How to prompt a model with tens or hundreds of thousands of tokens of input: why a bigger window is not a bigger memory, where to put documents and qu…

Read article →
ARTICLE · 14

Model-Specific Prompt Tricks, in depth: what really differs between Claude, OpenAI reasoning models and Gemini, and how to keep prompts portable

Model-specific prompting from first principles: the documented differences between Anthropic, OpenAI reasoning-model and Gemini guidance as of October…

Read article →
ARTICLE · 15

Multi-Agent Orchestration, in depth: briefs, handoffs and synthesis prompts that hold up

How to orchestrate several LLM agents: when splitting helps, topologies, orchestrator and worker briefs, handoff schemas, synthesis and verification p…

Read article →
ARTICLE · 16

Negative Prompting, in depth: why do-not instructions misfire, how to rewrite them, and how to enforce bans outside the prompt

A practical guide to negative instructions in LLM prompts: why negation is fragile for language models, positive reframing with worked rewrites, when …

Read article →
ARTICLE · 17

Agent Prompt Architecture in Depth

Prompting an agent rather than a single-turn model: the system prompt as specification, tool descriptions and parameter naming, choosing between tools…

Read article →
ARTICLE · 18

Chain-of-Thought Prompting, in Depth: Why Reasoning Tokens Help, How to Elicit and Parse Them, When They Hurt, and How to Measure the Gain

A practical guide to chain-of-thought prompting: why generating intermediate steps improves multi-step accuracy, zero-shot and few-shot patterns, answ…

Read article →
ARTICLE · 19

Chain-of-Verification Architecture, in depth: planning, factored verification and revision as a production pipeline

How Chain-of-Verification (CoVe) reduces hallucination: the four stages from the 2023 Meta AI paper, why joint, two-step, factored and factor+revise v…

Read article →
ARTICLE · 20

Context packing architecture

How to assemble the highest-value tokens into a fixed LLM context window: the anatomy of a context pack, salience scoring and reranking, token budgeti…

Read article →
ARTICLE · 21

Dynamic few-shot prompting

Deep-dive on dynamic few-shot prompting, where the demonstrations shown to a model are selected per request from an exemplar store rather than frozen …

Read article →
ARTICLE · 22

Prompt Evaluation Architecture in Depth

Evaluating prompts and LLM systems: representative eval sets, the three scorer families, LLM-as-judge bias and validation, pairwise comparison, sampli…

Read article →
ARTICLE · 23

Few-shot prompting architecture, in depth: designing, formatting, ordering, caching and evaluating a fixed example set

Few-shot prompting treated as a system: why in-context demonstrations work, what an example set must cover, inline tags versus multi-turn formatting, …

Read article →
ARTICLE · 24

LLM hallucination guardrails architecture

Deep-dive on layered hallucination guardrails for LLM systems: why fabrication is structural and trust asymmetric, grounding via retrieval, the cite-o…

Read article →
ARTICLE · 25

Least-to-most prompting architecture

Deep-dive on least-to-most prompting: decomposing a hard problem into an ordered easiest-to-hardest subproblem queue, passing each committed subanswer…

Read article →
ARTICLE · 26

Meta-prompting -- using an LLM to write and improve prompts

Deep-dive on meta-prompting: using an LLM to generate and refine prompts, the optimization loop (generate/evaluate/refine), the meta-prompt, automatic…

Read article →
ARTICLE · 27

Multimodal prompting architecture

Deep-dive on prompting with images: how pixels become tokens via resize and tiling, low vs high detail cost, image placement and interleaving, the cro…

Read article →
ARTICLE · 28

Output Formatting, in depth: choosing and specifying the shape of model output for people, parsers and renderers

How to decide what shape an LLM's output …

Read article →
ARTICLE · 29

Output Parsing Architecture in Depth: Turning Model Text into Trusted, Typed Data

How to build the parsing layer between an LLM and your code: classifying how a response ended, extracting payloads from free text, a bounded tolerant-…

Read article →
ARTICLE · 30

Prompt caching architecture

Deep-dive on LLM prompt caching: storing the deterministic KV attention state of a stable prompt prefix so later requests prefill only the new tail, t…

Read article →
ARTICLE · 31

Prompt compression architecture

Deep-dive on compressing LLM prompts at scale: segment inventory and per-class token budgets, relevance scoring, the extractive-abstractive-pruning ca…

Read article →
ARTICLE · 32

Prompt-injection defense architecture

Deep-dive on defending LLM applications against prompt injection: why no prompt can stop it, how to separate trusted instructions from untrusted conte…

Read article →
ARTICLE · 33

Prompt Pipeline Architecture in Depth

A 2500-word walkthrough of a production prompt pipeline: templates, variables, compiler, structured outputs, model router, validation, registry, eval,…

Read article →
ARTICLE · 34

Prompt registry architecture

Deep-dive on prompt registries: immutable prompt versions with variable schemas, eval-gated promotion through environment labels, cached runtime resol…

Read article →
ARTICLE · 35

Prompt Templates, in depth: typed contracts, strict rendering, token budgets, cache-friendly ordering and golden-render tests

How to design prompt templates as software: typed input contracts, strict and sandboxed rendering, the double-render injection bug, delimiting untrust…

Read article →
ARTICLE · 36

ReAct prompting

Deep-dive on ReAct prompting: the thought-action-observation loop, tool grounding for reduced hallucination, prompt format vs native tool calling, err…

Read article →
ARTICLE · 37

Reflexion architecture

Deep-dive on the Reflexion pattern: how an LLM agent improves within a single session through verbal reinforcement — writing natural-language self-cri…

Read article →
ARTICLE · 38

Role Prompting, in depth: what a role actually changes, what the evidence says about personas, writing role blocks as specifications, and testing them with an A/B harness

A practical guide to role prompting: how a role conditions a model, what research on personas in system prompts found about accuracy, the four things …

Read article →
ARTICLE · 39

Prompt routing architecture

Deep-dive on prompt routing: intent classifier, policy engine, LLM registry, selection, fallback ladder, quality gate, and A/B.

Read article →
ARTICLE · 40

Self-consistency -- sample many reasoning paths, vote

Deep-dive on self-consistency: sampling multiple chain-of-thought reasoning paths and taking the majority-vote answer, marginalizing over the reasonin…

Read article →
ARTICLE · 41

Semantic routing architecture

Deep-dive on semantic routing for LLM apps: embedding queries, labeled route centroids, cosine similarity and confidence thresholds, LLM fallback for …

Read article →
ARTICLE · 42

Skeleton-of-Thought prompting architecture

Deep-dive on Skeleton-of-Thought: a cheap skeleton call that lists an answer&a…

Read article →
ARTICLE · 43

Step-back prompting architecture

Deep-dive on step-back prompting: the two-stage abstraction-then-reasoning pipeline, using the derived principle as a retrieval key, self-verification…

Read article →
ARTICLE · 44

Structured Output Architecture for LLMs in Depth

A 2500-word walkthrough of structured output: schema, provider features (function calling, JSON mode), constrained decoding, validation, retry, stream…

Read article →
ARTICLE · 45

Tree of Thoughts architecture

Deep-dive on Tree of Thoughts prompting: thought decomposition, candidate generation, value vs vote state evaluation, BFS/DFS search with beam width a…

Read article →
ARTICLE · 46

Verifier Architecture, in depth: separating generation from checking, verifier types, calibrated thresholds, best-of-N, reject-and-retry and fallbacks

How to design an LLM verifier layer: why checking is easier than generating, programmatic checks, model graders and outcome versus process reward mode…

Read article →
ARTICLE · 47

Zero-Shot Prompting, in depth: Writing Instruction-Only Prompts as Specifications, Evaluating Them, and Knowing When to Escalate

Production zero-shot prompting: the five parts of an instruction-only prompt, label definitions and output contracts, zero-shot chain of thought, vali…

Read article →
ARTICLE · 48

Prompt A/B Testing in Production, in depth: assignment per user, exposure logging, judge metrics, sample size, the delta method and prompt-specific confounders

How to run online A/B tests of LLM prompts: where they sit after offline evals, randomizing by user or conversation, logging prompt hashes and exposur…

Read article →
ARTICLE · 49

Analogical Prompting, in depth: letting the model write its own exemplars, when it helps and how it fails

How analogical prompting works: the model recalls relevant and distinct problems with solutions before solving the target, as proposed in Large Langua…

Read article →
ARTICLE · 50

Anatomy of a Prompt, in depth: what the model actually receives, the components and their jobs, role authority, ordering and ablation

A first-principles breakdown of an LLM prompt: how chat templates flatten messages into tokens, the components a production prompt is built from and w…

Read article →
ARTICLE · 51

Prompt Chaining, in depth: when to split a task, step contracts and validation gates, error-compounding arithmetic, and a worked support-ticket chain

How to design, build and debug a prompt chain: when a chain beats a single prompt, the common topologies, step contracts and deterministic gates betwe…

Read article →
ARTICLE · 52

Prompt Compression with LLMLingua, in depth: perplexity-based token pruning, question-aware LongLLMLingua, the LLMLingua-2 classifier and when compression pays

How the LLMLingua family compresses prompts: why token information is uneven, Selective Context, LLMLingua's budget controller and iterative toke…

Read article →
ARTICLE · 53

Debate Prompting, in depth: stance prompts, rebuttal rounds, judge rubrics and the baseline debate has to beat

A practical guide to the prompts behind multi-agent debate: independent and stance-assigned openings, the round-N rebuttal prompt with anti-conformity…

Read article →
ARTICLE · 54

Debugging Broken Prompts, in depth: capture, reproduce, minimise, classify, fix one thing and lock it in

A systematic workflow for debugging LLM prompts that misbehave: capture the exact rendered request, measure the failure rate across runs, minimise the…

Read article →
ARTICLE · 55

Prompt Delimiters

Separate instructions from data using delimiters: XML tags, markdown headers, or triple quotes. Reduces injection risk and improves model comprehensio…

Read article →
ARTICLE · 56

Prompting for Code Generation, in depth: repository context, tests as the specification, diff contracts, the execute-and-repair loop and pass@k evaluation

How to prompt language models to write code that survives review: what context to assemble from the repository, using tests as the specification, choo…

Read article →
ARTICLE · 57

Domain-Specific Prompts, in depth: text-to-SQL prompts built from your schema, your definitions and your verified queries

How to prompt a model to write correct SQL against your own database: representing the schema, schema linking, a business glossary, retrieved verified…

Read article →
ARTICLE · 58

Prompt Engineering Team Structure, in depth: the work that exists, three operating models, roles, ownership in the repo and model migrations

How to organise people around prompts and LLM features: an inventory of the real work, centralised, embedded and hub-and-spoke models, a roles and res…

Read article →
ARTICLE · 59

Extraction Prompts, in depth: evidence quotes, absent versus unclear, normalising in code, long documents and per-field evaluation

How to write and run prompts that extract facts from documents reliably: field definitions that remove ambiguity, exact-quote evidence verified agains…

Read article →
ARTICLE · 60

Multilingual Prompting, in depth: pinning the output language, glossaries, locale formatting in code and per-language evaluation

How to build prompts that work across languages: choose the instruction language, pin the output language and register, protect product terms with a g…

Read article →
ARTICLE · 61

Prompt Design for Streaming UX, in depth: answer-first prompts, stream-safe formats, markers, partial JSON, buffered moderation and cancellation

How to design prompts and clients for streamed LLM output: why order matters once users read as tokens arrive, time-to-useful-token, answer-first prom…

Read article →
ARTICLE · 62

Temperature, Top-p, Top-k, in depth: what each knob does to the next-token distribution, why order matters, what providers allow, and how to tune settings per task

A practitioner's guide to sampling parameters: softmax and temperature worked through with real numbers, top-k, top-p and min-p trunc…

Read article →
ARTICLE · 63

Prompt Template Libraries

Reusable prompt templates, variables, composition, and library trade-offs. LangChain PromptTemplate, Guidance, Jinja2 for structured prompts that scal…

Read article →
ARTICLE · 64

Prompt Versioning, in depth: what a version contains, content-addressed releases, model pinning, eval-gated changes and rollback

How to version prompts properly: the full bundle a version must capture, content hashes as identity with aliases for release, change classes, a git an…

Read article →
ARTICLE · 65

Reflexion, in depth: building the loop, designing evaluators and reflections, and reading what the paper's numbers really show

A practical guide to implementing Reflexion, the pattern where an LLM agent retries a task with written lessons from its own failures: what the origin…

Read article →
ARTICLE · 66

RAG, in depth: chunking, hybrid retrieval, reranking, packing the prompt, evaluating each stage and the failure modes

Retrieval-augmented generation from first principles: why it works, structure-aware chunking with context headers, hybrid dense and lexical retrieval …

Read article →
ARTICLE · 67

Role Prompting, in depth: persona specs that survive long conversations, drift, simulated users and roleplay attacks

How to design, build and test long-lived personas for LLM products: writing a persona spec, separating persona from policy, why personas drift over ma…

Read article →
ARTICLE · 68

Self-Consistency, in depth: answer extraction, voting, early stopping and agreement as a confidence signal

Engineer self-consistency for production: why answer extraction and normalisation decide whether voting works, a provider-agnostic parallel implementa…

Read article →
ARTICLE · 69

Self-Refine, in depth: generate, critique and rewrite with one model, what the evidence shows, prompt design, a working loop and when not to use it

A practical guide to Self-Refine prompting: the generate, feedback and refine loop from Madaan et al. (2023), what the paper measured and where it fai…

Read article →
ARTICLE · 70

Structured Output, in depth: one JSON Schema contract, compiled for strict modes, validated twice and versioned

Treat LLM JSON output as a versioned contract: one schema source, compiled to the strict subsets that Claude and OpenAI enforce, stop-reason gating, f…

Read article →
ARTICLE · 71

Task Decomposition, in depth: where to cut a task, plans as validated graphs, and when not to split

How to decompose tasks for language models: seam tests for where to cut, static versus model-generated plans, a JSON plan validated as a DAG, a parall…

Read article →
ARTICLE · 72

Zero-Shot Prompting, in depth: why instruction-tuned models follow a bare instruction, scoring labels with log-probabilities, calibrating label bias and measuring prompt sensitivity

Zero-shot prompting from the measurement side: why instruction tuning makes instruction-only prompts work, generate-and-parse versus scoring each labe…

Read article →