Skip to content

Latest commit

 

History

History
381 lines (322 loc) · 48 KB

File metadata and controls

381 lines (322 loc) · 48 KB

📖 Glossary

Every term this repo uses, defined in one line, with a link to the page that treats it properly. It is a lookup table, not a study guide: use it to unblock yourself mid-question, then read the topic when you have time.

The acronym table below is for scanning when someone drops an initialism you half-recognise. The A to Z after it has the actual definitions.

Related: The AI Engineer 75 · Night-before cheat sheet · Study plans · Company pages · Role guides


🔤 Acronyms at a glance

Acronym Expands to Acronym Expands to
A2A Agent to agent ACL Access control list
ANN Approximate nearest neighbour ASR Automatic speech recognition
BPE Byte-pair encoding CFG Classifier-free guidance
CoT Chain of thought DPO Direct preference optimization
ECE Expected calibration error EMA Exponential moving average
FLOP Floating-point operation FSDP Fully sharded data parallel
GQA Grouped-query attention GRPO Group relative policy optimization
HBM High-bandwidth memory HITL Human in the loop
HNSW Hierarchical navigable small world ICL In-context learning
ITL Inter-token latency IVF Inverted file index
KV Key and value LLM Large language model
LoRA Low-rank adaptation MCP Model Context Protocol
MHA Multi-head attention MLE Maximum likelihood estimation
MoE Mixture of experts MQA Multi-query attention
MRR Mean reciprocal rank MSE Mean squared error
MTEB Massive Text Embedding Benchmark nDCG Normalized discounted cumulative gain
OCR Optical character recognition PEFT Parameter-efficient fine-tuning
PII Personally identifiable information PP Pipeline parallelism
PPO Proximal policy optimization PQ Product quantization
PSI Population stability index QLoRA Quantized LoRA
QPS Queries per second RAG Retrieval-augmented generation
RLAIF RL from AI feedback RLHF RL from human feedback
RLVR RL from verifiable rewards RMF Risk Management Framework (NIST AI)
ROC Receiver operating characteristic RoPE Rotary position embedding
RRF Reciprocal rank fusion SFT Supervised fine-tuning
SGD Stochastic gradient descent SLO Service level objective
SSE Server-sent events STT Speech to text
TP Tensor parallelism TPOT Time per output token
TTFT Time to first token TTS Text to speech
VAD Voice activity detection VAE Variational autoencoder
ViT Vision transformer VLM Vision-language model
WER Word error rate ZDR Zero data retention

A

  • A2A - the agent-to-agent protocol for calling a peer agent across a team or company boundary, built on Agent Cards and long-running tasks. 06 Agents
  • A/B test - an online experiment splitting users between variants to see which one moves the real product metric. 07 Evals
  • Accuracy - the share of predictions that are correct, and a liar on imbalanced data where predicting the majority class scores 99%. 01 ML foundations
  • ACL filtering - enforcing per-user permissions inside the retrieval query itself, never by asking the model to withhold results. 04 RAG
  • Adam - an optimizer with a per-parameter adaptive step size derived from moving averages of the gradient and its square. 01 ML foundations
  • AdamW - Adam with weight decay decoupled from the adaptive update, which is why it is the transformer default. 01 ML foundations
  • Agent - a model plus tools plus context, running in a loop with stop conditions, deciding for itself what to do next. 06 Agents
  • Agent Card - JSON metadata published at a well-known URI declaring an A2A agent's identity, skills, endpoint and auth requirements. 06 Agents
  • Agent Skills (SKILL.md) - a folder of procedural instructions and optional scripts loaded by three-stage progressive disclosure. 06 Agents
  • Agentic RAG - retrieval exposed as a tool the model calls in a loop, refining the query between calls instead of retrieving once. 04 RAG
  • ALiBi - positional handling that adds a linear distance penalty to attention scores instead of using position embeddings. 02 LLM fundamentals
  • Alignment - post-training that shapes behaviour, both helpfulness and refusals, into the weights themselves. 09 Safety
  • ANN - approximate nearest neighbour search, trading exact recall for sublinear query time once brute force stops scaling. 04 RAG
  • Arithmetic intensity - FLOPs performed per byte moved, the ratio that decides whether an operation is compute-bound or memory-bound. 08 Inference
  • Attention - each token emits a query, key and value, and mixes other positions' values weighted by scaled dot-product scores. 02 LLM fundamentals
  • AWQ - activation-aware 4-bit post-training quantization that scales the channels that matter most before rounding. 08 Inference

B

  • Barge-in - letting a user interrupt a speaking voice agent, which means cancelling TTS and generation and truncating state to what was actually heard. 10 Multimodal
  • Batch API - an asynchronous bulk endpoint priced at roughly half the synchronous rate, for work that tolerates delay. 08 Inference
  • BatchNorm - normalization across the batch dimension, unusable in transformers because of variable-length sequences and batch-size-1 decoding. 01 ML foundations
  • Beam search - decoding that keeps several partial sequences to maximise total likelihood, good for translation and ASR, bland and repetitive for open-ended text. 02 LLM fundamentals
  • Bias-variance decomposition - expected test error as bias squared plus variance plus irreducible noise, diagnosed from the train and validation gap. 01 ML foundations
  • Bi-encoder - embeds query and document independently, which is what makes precomputation and ANN search possible and what caps quality. 04 RAG
  • BM25 - the classic lexical ranking function, which nails IDs, SKUs, error codes and fresh jargon that dense retrieval misses. 04 RAG
  • BPE - byte-pair encoding, a tokenizer trained by repeatedly merging the most frequent adjacent pair until the vocab is full. 02 LLM fundamentals
  • Bradley-Terry - the pairwise preference model behind reward-model training and Arena-style Elo rankings. 05 Fine-tuning

C

  • C2PA / Content Credentials - a cryptographically signed provenance manifest attached to media, which strips on re-encode or screenshot. 10 Multimodal
  • Calibration - whether a predicted 0.8 actually means 80%, a property separate from ranking quality. 01 ML foundations
  • CaMeL - a hardened plan-then-execute design where the planner writes code from the trusted request alone and an interpreter enforces data-flow policy. 09 Safety
  • Canary eval - a scheduled eval run against the live production config so silent upstream changes page you instead of your users. 07 Evals
  • Catastrophic forgetting - general capability regressing after fine-tuning too hard on a narrow task, mitigated with lower LR, fewer epochs and mixed-in general data. 05 Fine-tuning
  • Causal mask - the upper-triangular minus-infinity mask applied before softmax so a position cannot attend to later tokens. 02 LLM fundamentals
  • CFG - classifier-free guidance, extrapolating between conditional and unconditional denoising to trade diversity for prompt adherence. 10 Multimodal
  • Chain of thought - prompting the model to write intermediate reasoning before the answer, valuable on multi-step problems and wasted latency on lookups. 03 Prompting
  • Chat template - the exact special-token markup a model was post-trained with, and the single most common silent killer of fine-tunes. 05 Fine-tuning
  • Chinchilla - the scaling result that compute-optimal training grows parameters and tokens together, roughly 20 tokens per parameter. 02 LLM fundamentals
  • Chunked prefill - splitting a long prefill into pieces mixed into decode batches so one huge prompt does not stall every other user's stream. 08 Inference
  • Chunking - splitting documents into retrievable units, typically 256 to 1024 tokens with 10 to 20% overlap. 04 RAG
  • Circuit breaker - a per-provider trip that stops sending traffic to a dependency that is already failing. 08 Inference
  • CLIP - contrastive image-text pretraining that lands both modalities in one embedding space, which is why zero-shot classification and cross-modal retrieval work. 10 Multimodal
  • ColBERT / late interaction - per-token embeddings scored by summed max similarity, sitting between bi-encoder speed and cross-encoder accuracy. 04 RAG
  • ColPali - screenshot-based document retrieval that embeds page images directly, so nothing is lost in parsing. 10 Multimodal
  • Compaction - summarising old turns while keeping recent ones verbatim, so an agent's context stays usable over a long trajectory. 06 Agents
  • Constitutional AI - alignment in which the model critiques and revises its own output against a written set of principles. 05 Fine-tuning
  • Constrained decoding - compiling a schema into a grammar and masking invalid tokens in the logits, which guarantees syntax but not semantics. 03 Prompting
  • Contamination - eval data present in the pretraining corpus, so the score measures memorisation rather than capability. 07 Evals
  • Context engineering - managing the whole token budget across turns (tools, retrieved docs, memory, history), not just optimising one string. 03 Prompting
  • Context rot - quality decay as stale errors, bloated tool output and irrelevant retrievals accumulate, well below the hard context limit. 06 Agents
  • Contextual retrieval - prepending a short LLM-written blurb situating each chunk in its document before embedding it. 04 RAG
  • Continued pretraining - more next-token training on domain data, the right move when the domain language itself is foreign rather than just the facts. 05 Fine-tuning
  • Continuous batching - composing the serving batch per decode step so finished requests leave and queued ones join mid-flight. 08 Inference
  • Cosine similarity - direction-only similarity that ignores magnitude, the default for comparing text embeddings. 01 ML foundations
  • Covariate shift - the input distribution moves while the input-to-label relationship holds, as opposed to concept drift where the relationship itself changes. 01 ML foundations
  • Cross-encoder - scores query and document jointly through one transformer with full attention, accurate but one forward pass per candidate. 04 RAG
  • Cross-entropy - the maximum-likelihood loss for categorical outputs, whose gradient through softmax is simply predicted minus actual. 01 ML foundations
  • Cross-validation - rotating held-out folds to use a small dataset efficiently, replaced by temporal splits for anything time-dependent. 01 ML foundations

D

  • Data leakage - training information reaching the evaluation, in forms as subtle as fitting a scaler before splitting or the same user in both sets. 01 ML foundations
  • Decode - the sequential phase that emits one token at a time, memory-bandwidth-bound because every weight must be streamed per token. 08 Inference
  • Decoder-only - the causal next-token architecture behind every mainstream LLM, where every position yields a training signal. 02 LLM fundamentals
  • Defence in depth - stacking input, output and action controls because no single guardrail holds against an adaptive attacker. 09 Safety
  • Diffusion - generation by learning to reverse a gradual noising process, predicting the noise added at each step. 10 Multimodal
  • Distillation - training a small model on a large model's outputs or logits, the standard way to compress reasoning behaviour into a cheap model. 05 Fine-tuning
  • DiT - the diffusion transformer backbone that replaced the U-Net in recent image models. 10 Multimodal
  • Double descent - the observation that heavily overparameterized networks generalise better past the interpolation threshold, breaking the classical U-curve. 01 ML foundations
  • DPO - direct preference optimization, training on preference pairs with a classification loss and no reward model or RL loop. 05 Fine-tuning
  • Drift - the family of production failures where model version, input distribution, or cost and latency move under you without a deploy. 07 Evals
  • Dropout - randomly zeroing activations at train time as an implicit ensemble, usually set to zero in large-scale pretraining. 01 ML foundations
  • Dual-LLM pattern - a privileged model that plans and calls tools but never reads untrusted content, paired with a quarantined model that reads it and returns typed variables. 09 Safety

E

  • Early stopping - halting training when validation loss stops improving, cheap and close to L2 in effect. 01 ML foundations
  • ECE - expected calibration error, the summary number from a reliability diagram, fixed post hoc with temperature scaling. 01 ML foundations
  • Elo / Arena - pairwise human preference ranking, hard to contaminate but measuring preference and style rather than task correctness. 07 Evals
  • Embedding - a vector whose geometry encodes meaning, compared with whatever similarity the model was trained with. 01 ML foundations
  • Emergent abilities - apparent sharp capability jumps with scale, partly a measurement artifact of discontinuous metrics. 02 LLM fundamentals
  • Encoder-decoder - a bidirectional encoder feeding a cross-attending decoder, still strong for fixed input-to-output transforms like translation and ASR. 02 LLM fundamentals
  • Encoder-only - bidirectional masked-LM models such as BERT, which cannot generate and now live on as embedding and reranker models. 02 LLM fundamentals
  • Error analysis - reading 50 to 100 failing traces, clustering the failure descriptions, and fixing the biggest cluster first. 07 Evals
  • Eval set - the versioned dataset that encodes what good means for your product, and the asset that survives every model swap. 07 Evals
  • Exfiltration channel - any path by which data can leave the system, the third leg of the lethal trifecta. 09 Safety
  • Exponential backoff with jitter - retry spacing that randomises the delay so clients do not synchronise into a retry storm. 08 Inference

F

  • Faithfulness - whether every claim in a generated answer is supported by the retrieved context, usually judged claim by claim. 07 Evals
  • Few-shot - putting worked examples in the prompt to anchor format and sharpen fuzzy decision boundaries. 03 Prompting
  • FID - a distribution-level realism score for generated images, meaningless for judging any single image. 10 Multimodal
  • Fine-tuning - further training that changes form, style and narrow skill, and the wrong tool for injecting facts. 05 Fine-tuning
  • FlashAttention - an IO-aware exact attention implementation that tiles into on-chip SRAM and never materialises the full score matrix. 02 LLM fundamentals
  • FP8 / INT8 - eight-bit formats that shrink weights and, when activations are quantized too, accelerate the matmuls on tensor cores. 08 Inference
  • FSDP / ZeRO - sharding optimizer states, gradients and parameters across GPUs so a model too large for one device can still train. 05 Fine-tuning

G

  • GCG - the gradient-searched adversarial suffix attack, notable because the strings transfer across models. 09 Safety
  • GGUF - the quantized model file format used by the llama.cpp and Ollama local-inference ecosystem. 08 Inference
  • Golden set - a labelled set of real, hard cases you gate prompt and model changes on. 11 System design
  • Goodhart's law - once a measure becomes a target it stops being a good measure, which is what happens to every headline benchmark. 07 Evals
  • Goodput - throughput that actually meets your latency SLO, the only throughput number worth reporting. 08 Inference
  • GPTQ - Hessian-based 4-bit post-training weight quantization, one of the two standard methods alongside AWQ. 08 Inference
  • GQA - grouped-query attention, where groups of query heads share one KV head, cutting cache size at near-MHA quality. 02 LLM fundamentals
  • Gradient checkpointing - recomputing activations in the backward pass instead of storing them, roughly 30% slower for a large memory win. 05 Fine-tuning
  • Gradient clipping - capping the global gradient norm, conventionally at 1.0, so a loss spike does not destroy the run. 01 ML foundations
  • GraphRAG - building an entity and relationship graph at index time, worth the cost for relational or corpus-wide questions, not for factoid lookup. 04 RAG
  • Grounding - constraining the model to answer from supplied sources and cite them, the first line of defence against confident fabrication. 09 Safety
  • GRPO - group relative policy optimization, scoring each response against its group mean instead of training a value network. 05 Fine-tuning
  • Guardrail metric - a must-never-regress number such as PII leakage or jailbreak rate, gated as a binary, never traded for quality. 07 Evals
  • Guardrails - the input, output and action checks around a model call: classifiers, moderation, schema validation and approval gates. 09 Safety

H

  • Hallucination - confident fabrication, structural because the training objective rewards plausible next tokens rather than truth. 02 LLM fundamentals
  • Handoff - transferring conversation control to a specialist agent, the multi-agent pattern that suits distinct domains like support triage. 06 Agents
  • HBM - the GPU's high-bandwidth memory, whose bandwidth sets the hard ceiling on batch-1 decode speed. 08 Inference
  • HITL - human in the loop, the review queue and approval gate for low-confidence outputs and high-risk actions. 11 System design
  • HNSW - a multi-layer navigable small-world graph index, the standard high-recall ANN structure, tuned with M and ef_search. 04 RAG
  • HumanEval - a set of 164 Python problems scored with pass@k, now saturated and small. 07 Evals
  • Hybrid search - running lexical and dense retrieval together and fusing their rankings, the fix for queries containing IDs or jargon. 04 RAG
  • HyDE - having the model write a hypothetical answer and embedding that instead of the question, since answers sit closer to documents. 04 RAG

I

  • Idempotency key - a client-supplied token that makes a retried side-effectful call safe to repeat. 08 Inference
  • In-context learning - inferring the task from demonstrations in the prompt, with no weight update. 03 Prompting
  • Indirect prompt injection - instructions planted in content your app processes on someone else's behalf, dangerous because the victim never sees the attack. 09 Safety
  • InfoNCE - the contrastive loss behind CLIP and modern sentence embedders, where the positive pair must beat in-batch negatives. 01 ML foundations
  • Instruction hierarchy - training-time privileging of system over user over tool content, which lowers attack success rates but is probabilistic, not enforced. 09 Safety
  • Interleaving - mixing two rankers' results into one list to reach significance with far less traffic than an A/B test. 07 Evals
  • IVF - an inverted-file ANN index that clusters vectors into cells and probes only the nearest ones. 04 RAG

J

  • Jailbreak - an attack on the model's safety training to elicit forbidden content, a different attacker and owner from prompt injection. 09 Safety
  • JSON mode - the model is instructed to emit JSON, so you still validate and retry, unlike constrained decoding. 03 Prompting
  • Judge (LLM-as-judge) - grading open-ended output with another model and a rubric, reliable only after you measure its agreement with humans. 07 Evals

K

  • KL penalty - the term keeping an RL-tuned policy near its frozen reference, and the thing whose removal invites reward hacking. 05 Fine-tuning
  • KTO - preference tuning that works on binary good and bad labels, so you do not need paired comparisons. 05 Fine-tuning
  • KV cache - stored keys and values for every past token, turning quadratic recompute into linear lookups at a large memory cost. 02 LLM fundamentals
  • KV cache quantization - storing keys and values in FP8 or INT8, a lever that directly buys batch size and therefore throughput. 08 Inference

L

  • Late chunking - embedding the whole document with a long-context embedder first, then pooling token embeddings per chunk. 04 RAG
  • Latent diffusion - running the diffusion process in a compressed VAE latent space rather than on pixels, which is what makes it tractable. 10 Multimodal
  • LayerNorm - normalization across features within each token, batch-independent and therefore fine at batch size 1. 01 ML foundations
  • Lethal trifecta - private data access plus untrusted content plus an exfiltration channel, the combination you design around by removing one leg. 09 Safety
  • LIMA - the result that roughly 1,000 meticulously curated examples can align a large base model, quality over quantity. 05 Fine-tuning
  • Llama Guard - a safeguard model that classifies content against a hazard taxonomy, used as an input or output filter. 09 Safety
  • Logit masking - the mechanism constrained decoding uses, setting invalid tokens to minus infinity before sampling. 03 Prompting
  • Logprobs - per-token log probabilities from the API, your tool for confidence scoring, classification and perplexity evals. 02 LLM fundamentals
  • LoRA - a low-rank adapter update, W' = W + (alpha/r)BA, with base weights frozen and trainable parameters at roughly 0.1 to 1% of the model. 05 Fine-tuning
  • Lost in the middle - the U-shaped finding that models attend best to the start and end of a long context, so key material belongs at the edges. 03 Prompting

M

  • Many-shot jailbreak - hundreds of faux dialogue turns exploiting in-context learning in a long window to override safety training. 09 Safety
  • Matryoshka embeddings - vectors trained so their prefixes are valid embeddings, letting you truncate for a cheap first pass and refine later. 04 RAG
  • MCP - the Model Context Protocol, an open standard turning N times M bespoke integrations into N plus M, with tools, resources and prompts as server primitives. 06 Agents
  • Memorisation - models reproducing training data verbatim, which is why training and fine-tuning corpora need dedup and PII scrubbing. 09 Safety
  • Metadata filtering - restricting an ANN search by tenant, date or ACL, which has to happen inside the index traversal rather than after top-k. 04 RAG
  • MHA - multi-head attention, splitting the model dimension into parallel heads so different relations can be attended to at once. 02 LLM fundamentals
  • Mixed precision - training with bf16 weights and gradients alongside fp32 master weights and optimizer states, roughly 16 bytes per parameter with Adam. 05 Fine-tuning
  • MLA - DeepSeek's attention variant that low-rank-compresses the KV cache, going further than GQA on memory. 02 LLM fundamentals
  • MMLU - a 57-subject multiple-choice knowledge benchmark, saturated at the frontier and widely contaminated. 07 Evals
  • Modality gap - the observation that image and text embeddings occupy separated cones inside CLIP's shared space. 10 Multimodal
  • Mode collapse - synthetic training data amplifying the teacher model's stylistic tics until output diversity dies. 05 Fine-tuning
  • Model card / system card - documentation of intended use, evals and limitations, for the model and for your deployed system respectively. 09 Safety
  • Model gateway - the single choke point for every LLM call, handling auth, quotas, retries, provider failover and usage metering. 11 System design
  • MoE - mixture of experts, routed expert MLPs giving large total parameters with a small active count per token, at the cost of holding every expert in memory. 02 LLM fundamentals
  • Momentum - keeping a moving average of gradients to damp oscillation across ravines and accelerate consistent directions. 01 ML foundations
  • MQA - multi-query attention, where all query heads share a single KV head, shrinking the cache at some quality cost. 02 LLM fundamentals
  • MRR - mean reciprocal rank, scoring how high the first relevant result appeared. 04 RAG
  • MTEB - the standard embedding benchmark, useful for shortlisting and useless as final proof on your own domain. 04 RAG
  • Multi-agent - an orchestrator delegating to workers, which wins on parallel read-heavy work and hurts on write-heavy shared state. 06 Agents
  • Multimodal RAG - retrieval over visually rich documents, usually by captioning figures at ingestion or embedding page images directly. 10 Multimodal

N

  • nDCG - a graded-relevance ranking metric that scores the whole result list, not just the first hit. 04 RAG
  • NF4 - the 4-bit data type QLoRA quantizes the frozen base to, designed for normally distributed weights. 05 Fine-tuning
  • Nucleus sampling (top-p) - keeping the smallest set of tokens whose cumulative probability reaches p, so the cutoff adapts to model confidence. 02 LLM fundamentals

O

  • Observability - tracing every model call, tool call and retrieval as spans with tokens, latency and cost attached, because you cannot debug what you did not record. 07 Evals
  • OCR-free extraction - sending the page image straight to a VLM instead of running a text-recognition pipeline first. 10 Multimodal
  • Online eval - measuring on live traffic through A/B tests, interleaving and implicit signals, which confirms what offline evals only predict. 07 Evals
  • OpenTelemetry GenAI conventions - the standard attribute names for LLM spans, so traces stay portable across backends. 07 Evals
  • Orchestrator-worker - the subagent pattern whose real payoff is context isolation: a worker burns tokens and returns a short summary. 06 Agents
  • ORPO - preference optimisation folded into the SFT loss, with no separate reference model. 05 Fine-tuning
  • Over-refusal - refusing benign requests, the failure mode that makes a model trivially safe and commercially useless. 09 Safety
  • OWASP Top 10 for LLM Applications - the shared vocabulary of AI security reviews, from prompt injection through unbounded consumption. 09 Safety

P

  • PagedAttention - managing the KV cache in fixed-size blocks through a block table, virtual memory for the cache, which kills fragmentation. 08 Inference
  • Parent-document retrieval - matching on small precise chunks and returning the larger parent for context, also called small-to-big. 04 RAG
  • pass@k - the probability that at least one of k samples solves the problem, estimated without bias from n samples and c correct. 07 Evals
  • pass^k - the probability that all k trials succeed, the consistency metric that matters for agents and that pass@k hides. 07 Evals
  • PEFT - parameter-efficient fine-tuning, training a small set of added parameters so no optimizer state is needed for frozen weights. 05 Fine-tuning
  • Pipeline parallelism - sharding a model by layer ranges across devices, tolerant of slow interconnect but adding latency and bubbles. 08 Inference
  • Position bias - a judge favouring whichever response it saw first, mitigated by running both orders and keeping consistent verdicts. 07 Evals
  • Position interpolation - rescaling positions back inside the trained range, plus a short fine-tune, to extend usable context. 02 LLM fundamentals
  • PPO - the RL algorithm in the classic RLHF recipe, optimising the policy against a reward model under a KL penalty. 05 Fine-tuning
  • PQ - product quantization, compressing vectors into subspace codebook codes for large memory savings at some recall cost. 04 RAG
  • PR-AUC - the honest ranking metric when positives are rare, with the prevalence rather than 0.5 as its baseline. 01 ML foundations
  • Precision and recall - the share of predicted positives that are correct, and the share of actual positives that were found. 01 ML foundations
  • Prefill - processing the whole prompt in one parallel pass, compute-bound, and the phase that sets time to first token. 08 Inference
  • Pre-norm - placing the norm inside the residual branch before each sublayer, which keeps the residual stream a clean identity path. 02 LLM fundamentals
  • Prompt caching - reusing the KV cache of a shared prefix across requests, which is why prompts should run stable content first and volatile content last. 03 Prompting
  • Prompt injection - attacking the application through content the model reads, unsolved because the context window has no privilege separation. 09 Safety
  • PSI - population stability index, a drift measure comparing a current input window against a reference one. 01 ML foundations

Q

  • QLoRA - a 4-bit NF4 frozen base with double quantization and paged optimizers, with LoRA trained on top in bf16. 05 Fine-tuning
  • Quantization - lowering weight or activation precision to cut memory and bandwidth, roughly free at 8-bit and measurable at 4-bit. 08 Inference
  • Query rewriting - turning a conversational follow-up into a standalone search query, non-negotiable for multi-turn RAG. 04 RAG

R

  • RAG - retrieval-augmented generation, supplying fresh or private knowledge at query time instead of baking it into weights. 04 RAG
  • RAGAS - a framework packaging the standard RAG metrics, faithfulness and answer relevance among them. 07 Evals
  • ReAct - interleaving thought, action and observation, now largely absorbed into native tool-calling loops. 06 Agents
  • Reasoning model - a model RL-trained to spend test-time compute on long chains of thought, buying accuracy on hard problems at higher latency and cost. 02 LLM fundamentals
  • Recall@k - whether the gold chunk made the top k, the retrieval metric that gates everything downstream. 04 RAG
  • Reflection - generate, critique, revise, which works when verification is grounded in an external signal and stalls after one or two rounds. 06 Agents
  • Regression test - running the eval suite on every prompt, model or retrieval change and blocking the merge on a regression. 07 Evals
  • Reliability diagram - a plot of predicted confidence against observed accuracy, the visual form of calibration. 01 ML foundations
  • Reranking - a second, more expensive pass that reorders a wide candidate list, typically retrieve 100 to 200 then keep 5 to 20. 04 RAG
  • Residual stream - the running hidden state that attention and MLP blocks read from and write into, the model's workspace. 02 LLM fundamentals
  • Retrieval miss vs generation miss - the first triage question for any bad RAG answer: were the right chunks fetched, or were they fetched and ignored? 04 RAG
  • Reward hacking - a policy exploiting the reward model rather than improving, showing up as sycophancy and confident bloat. 05 Fine-tuning
  • Reward model - a model trained on human preference pairs to score responses during RLHF. 05 Fine-tuning
  • RLAIF - reinforcement learning from AI feedback, replacing most human preference labels with model-generated ones guided by principles. 05 Fine-tuning
  • RLHF - the SFT, reward model and PPO pipeline that turned raw completion models into steerable assistants. 05 Fine-tuning
  • RLVR - reinforcement learning on verifiable rewards such as passing tests or correct answers, which is what trains reasoning models. 05 Fine-tuning
  • RMSNorm - LayerNorm without mean-centring, just a rescale by the root mean square, cheaper and equally effective. 01 ML foundations
  • ROC-AUC - the probability a random positive outranks a random negative, insensitive to class imbalance and therefore misleading on rare positives. 01 ML foundations
  • RoPE - rotary position embedding, rotating query and key pairs so their dot product depends only on relative position. 02 LLM fundamentals
  • Router - the component that sends each request to the right model tier, usually the single biggest cost lever in the system. 11 System design
  • RRF - reciprocal rank fusion, combining rankings by rank because BM25 scores and cosine similarities live on incomparable scales. 04 RAG
  • Rug pull - a third-party tool server changing its descriptions after you approved them, one of the MCP supply-chain attack classes. 09 Safety

S

  • Safetensors - the data-only weight format that cannot execute code on load, unlike pickle-based checkpoints. 09 Safety
  • Sandboxing - running model-generated code in an isolated container with no network egress, resource limits and a throwaway filesystem. 09 Safety
  • Scaling laws - power-law relationships between compute, data, parameters and loss, and the reason deployment-optimal differs from compute-optimal. 02 LLM fundamentals
  • Self-consistency - sampling k reasoning paths above temperature zero and majority-voting the final answer, at k times the cost. 03 Prompting
  • Self-preference bias - judges preferring output from their own model family, mitigated by judging with a different family. 07 Evals
  • Semantic caching - serving a stored answer for a semantically similar query, which needs care around ACLs and freshness. 08 Inference
  • SentencePiece - a language-agnostic tokenizer library that treats raw text as a stream and needs no pre-tokenization. 02 LLM fundamentals
  • SFT - supervised fine-tuning on prompt and response pairs rendered through the chat template, with loss masked to response tokens. 05 Fine-tuning
  • SigLIP - the sigmoid-loss successor to CLIP, now a common vision backbone for VLMs. 10 Multimodal
  • SLO - the latency or quality objective you size capacity against, for example P99 TTFT under 800 ms. 08 Inference
  • Speculative decoding - a cheap drafter proposes tokens the target model verifies in parallel, speeding up decode without changing the output distribution. 08 Inference
  • SSE - server-sent events, the transport that streams token deltas so perceived latency is TTFT rather than full completion time. 08 Inference
  • Structured output - a response constrained to a schema and validated deterministically, the highest-leverage guardrail available. 03 Prompting
  • Subagent - a worker with its own clean context window that returns a distilled summary to the orchestrator. 06 Agents
  • SWE-bench - resolving real GitHub issues, still discriminative but sensitive to the agent harness around the model. 07 Evals
  • System prompt leakage - assume it happens, so never put secrets or unenforced authorisation logic in it. 09 Safety

T

  • Temperature - the softmax rescaling knob, where zero is greedy argmax and above one flattens the distribution. 02 LLM fundamentals
  • Temperature scaling - post-hoc calibration that divides logits by a single fitted scalar without changing the ranking. 01 ML foundations
  • Tensor parallelism - sharding every layer's matrices across GPUs with an all-reduce per layer, which cuts latency but needs fast interconnect. 08 Inference
  • Token - the unit the model actually sees, a subword ID rather than a character or a word, which is why letter counting and arithmetic fail. 02 LLM fundamentals
  • Tool call - a structured message naming a tool and its JSON arguments, which your runtime executes, never the model. 06 Agents
  • Tool description - the prose the model reads to decide when to call a tool, effectively a prompt and worth writing like one. 06 Agents
  • Tool poisoning - malicious instructions hidden in a tool description that lands in your model's context. 09 Safety
  • Top-k sampling - keeping only the k highest-probability tokens before renormalizing and sampling. 02 LLM fundamentals
  • TPOT / ITL - time per output token after the first, driven by memory bandwidth and batch contention. 08 Inference
  • Trace and span - one request as a tree, with a span per model call, tool invocation, retrieval and guardrail check. 07 Evals
  • Trajectory eval - grading the path an agent took (tool choice, argument correctness, step efficiency), which diagnoses why outcomes failed. 07 Evals
  • TTFT - time to first token, dominated by queueing and prefill, and the number streaming UX is built around. 08 Inference
  • TTS - text to speech, classically an acoustic model plus vocoder, now usually a neural codec language model. 10 Multimodal

V

  • VAE - the autoencoder that compresses images into the latent space diffusion operates in, and decodes the result back to pixels. 10 Multimodal
  • Verbosity bias - longer answers scoring higher regardless of quality, mitigated by a rubric that penalises padding and by reporting length. 07 Evals
  • Verifiable reward - a programmatic correctness check such as a unit test or an answer checker, which is why RL for reasoning works on maths and code. 05 Fine-tuning
  • ViT - vision transformer, an image cut into fixed-size patches and run through a standard transformer. 10 Multimodal
  • vLLM - the default open-source serving stack, home of PagedAttention and continuous batching. 08 Inference
  • VLM - vision-language model, a vision encoder plus a projector plus an LLM that treats patch embeddings as ordinary tokens. 10 Multimodal

W

  • Warmup - the short linear learning-rate ramp that stops deep transformers diverging while Adam's second-moment estimate is still garbage. 01 ML foundations
  • Weight decay - shrinking weights toward zero, applied directly to the weights in AdamW rather than through the gradient. 01 ML foundations
  • WER - word error rate, the standard accuracy metric for speech recognition. 10 Multimodal
  • Workflow - LLMs and tools orchestrated through predefined code paths, which is what you should build whenever you can draw the flowchart. 06 Agents

Y

  • YaRN - context extension that interpolates RoPE frequency bands unevenly, preserving local resolution with much less fine-tuning. 02 LLM fundamentals

Z

  • ZDR - zero data retention, a contract term removing even the vendor's short abuse-monitoring retention window. 09 Safety
  • Zero-shot - asking the model to do the task with instructions only, no examples, which works well on instruction-tuned models. 03 Prompting

Where to go next

If you want Go to
The 75 highest-signal items in checkbox form AI-ENGINEER-75.md
One evening of revision before the interview CHEATSHEET.md
A structured 1, 4 or 8-week plan STUDY_PLAN.md
Questions and loop maps for a specific company 14-company-interview-questions
A study map calibrated to your job title 15-role-guides
Papers, courses and blogs worth the time resources

Missing a term? Corrections and additions are welcome, see CONTRIBUTING.md.