Resource

AI Glossary

Plain-English definitions for the vocabulary of AI systems, automation, and operational verification. Written for leaders evaluating or deploying AI in their organization, not for engineers.

71
Terms defined
7
Domains covered
GitHub
Source of truth

Vault and knowledge structure

Sprawl

The uncontrolled, inconsistent growth of a folder or file structure over time. It shows up as duplicate homes for the same kind of content, naming conventions that drift from folder to folder, and structure that no longer matches whatever documentation was meant to govern it. Sprawl is often confused with simple depth, but the two are different problems: a deeply nested structure can be perfectly clean if every level is deliberate and consistent, and a shallow structure can still be sprawling if it grew ad hoc. The tell for sprawl is not how many levels deep something is, it's whether new content keeps landing in the right place without anyone having to think about it.

De-stuttering

The cleanup pass that removes the repeated part of a stuttering name, so the prefix is left to carry that meaning by itself. The discipline worth keeping is that de-stuttering is run as its own pass, separately from the migration that created the stutter. A prefix migration is a single mechanical rule applied everywhere, which is exactly what makes it verifiable and reversible. Deciding which remaining words are redundant and which are load-bearing is a judgment call made file by file. Running both at once turns one auditable rule into two entangled ones, and leaves you unable to tell which change caused a given breakage.

Stuttering

When a name repeats itself because a prefix and the words after it carry the same meaning. A file named acme_acme-vendor-list stutters: the prefix already says Acme. It almost always appears after a naming convention is introduced or changed, because the older names spelled out what the new prefix now encodes, and a mechanical rename adds the prefix without removing the spelled-out version. On its own it looks cosmetic. At scale it is not: it makes every reference longer, it breaks autocomplete because typing the prefix stops narrowing anything, and it quietly teaches everyone that the prefix does not mean what it claims to mean.

File and content operations

Diff (verb: to diff / diffed)

To compare two versions of something (two files, two folders, two document drafts) item-by-item and identify exactly what's only in one, only in the other, or changed between them. Same root as a software "diff" in version control. Useful any time two things are supposed to be identical, such as a backup, a synced folder, or a migrated copy, and you need to actually prove that rather than assume it.

Canonical / Sole canonical

The one version of something treated as the authoritative source when more than one copy exists. Calling something the "sole canonical copy" is a deliberate statement that it's the only version anyone should read from or edit going forward; everything else is a mirror, backup, export, or derivative that should defer to it, never the other way around. The term shows up constantly in any system with duplicated or synced data (files, folders, databases, documentation) because without naming one copy canonical, it's ambiguous which version is "true" when two versions disagree. Canonical applies to containers as much as to individual files: which account, workspace, or system is the one of record. That layer is the easier one to get wrong, because two accounts holding identically named folders both look correct until you check which one you are actually in. Naming the canonical container is worth doing before naming canonical copies inside it.

Atomic replacement

With atomic replacement, a reader sees either the complete old file or the complete new file. It avoids exposing a half-written file. It does not, by itself, protect against another editor racing to change the same file or make several file changes succeed as one unit.

Compare-and-write

Before replacing a document, check that nobody has changed it since you last saw it. A true compare-and-write guarantee keeps that check and the update together. If they are separate steps, another editor can still slip in between them.

Page chrome

Page chrome is the material around the main content: page numbers, headers, footers, menus, and banners. Some of it is decoration; some carries context or makes citations possible. When a tool removes it, decide based on the document's purpose rather than assuming all surrounding material is disposable.

Recovery copy

A recovery copy preserves what a file contained before a change. It gives you something concrete to inspect or restore after a mistake. Keeping the copy is one part of a recovery plan; the process for deciding when and how to restore it is another.

Text layer

A PDF text layer is computer-readable text stored behind the visible page. It lets you select, search, and copy words; a scan may contain only an image until text recognition is added. Check for a usable text layer before choosing how to process or verify a PDF.

Verification and diagnostics

HTTP status code

The three digit code a web server returns with every response. 200 means it responded, 404 means the resource was not found, 403 means access was refused, 500 means the server itself failed. Useful as a first signal and insufficient as a final answer, because a server can return 200 while sending something entirely different from what was requested.

Content type

The declared format of what a server actually sent, for example an image or a web page. When you are verifying that something worked, content type is the more meaningful check, because it reveals cases where the request succeeded but the response was the wrong thing.

False-positive 200

A request that returns a success code while delivering the wrong thing entirely. In web systems, a server can answer 200 OK while actually serving an error page, a login redirect, or a placeholder. Anything checking only the status code sees success. This is the concrete, everyday version of a broader principle: a green signal is evidence that something responded, not evidence that it responded correctly. Always verify the content, not just the code.

Verification halo

Confirming the part of a system you touched, then treating the entire system as verified. The check that was performed was real, which is what makes this so easy to miss. It just covered a narrower scope than the confidence it produced. Confirming that a change was saved is not the same as confirming the change had an effect, and the gap between those two is where a surprising number of production failures live.

Silent success

An automated process that runs on schedule, reports success, and accomplishes nothing. The job starts, the log turns green, and no actual work happens. It is one of the most common and most expensive failures in automated systems, because every dashboard says everything is fine. The cause is almost always that the system measures execution rather than output: it checks whether the job ran, not whether the job did anything. The fix is to monitor results, not activity.

Control test

Running a case with a known outcome alongside the case you are actually investigating, so that a failure is distinguishable from a broken test. Without a control, an inconclusive result and a negative result look identical. It is a basic experimental discipline that is routinely skipped in technical troubleshooting, and skipping it is how teams confidently reach the wrong conclusion.

Cache-buster

Adding a unique value to a request to force a fresh response rather than a stored copy. Used when diagnosing whether you are looking at current state or at something a cache is holding onto. A useful habit before concluding that a fix did not work.

Ground truth

Ground truth is the original evidence you use to judge a result. For a converted document, that usually means checking the visible source page rather than comparing two converted copies. Two copies can agree with each other and still share the same mistake.

Anchor check

After converting a document, compare a few known points with the original: an important number, a difficult table entry, a figure label, something near the beginning, and the last meaningful line. Checking the end helps reveal text that was quietly cut off. A spot check is useful, but it should be paired with a check for missing material between those points.

Content hash

A digital fingerprint calculated from a file. Comparing fingerprints helps check whether two copies have the same contents or whether a file changed. A fingerprint alone does not tell you who created the file or whether its contents are trustworthy.

Coverage vs accuracy

Accuracy asks whether the parts you checked are correct. Coverage asks whether all the needed parts are present. A document conversion can pass several spot checks and still omit entire sections, so a good review checks both correctness and completeness.

Endpoint over docs

A working rule for integrating with fast-moving services: where published documentation and the live endpoint disagree, believe the endpoint. Documentation on rapidly shipping products is routinely behind the deployed reality. The cheapest way through is to send the smallest possible request and read the rejection carefully, because a well-built API usually lists its accepted values in the error. That makes a failed call a better specification than the guide, and it turns an afternoon of guessing into a few minutes of probing.

Free-baseline diff

Run a lower-cost or free method as a comparison point before paying for a more advanced one. Compare simple signals such as page, word, or number counts to spot unexplained omissions. A difference is a reason to investigate, not automatic proof that either result is wrong.

Homoglyph

A homoglyph is a character that looks like another character, such as the digit zero and the letter O. These mix-ups may be easy to infer in a sentence but dangerous in an account number, invoice ID, or other exact identifier. Check important identifiers against the original image or source.

Pre-registered prediction

Write down what you expect a test to show, and how you will judge it, before running the test. That makes a surprising result useful evidence rather than something to explain away afterward. The practice helps a team learn from both success and failure.

Regex gap

A pattern used to find or extract something is written broadly enough to match the common cases in a dataset but not a format that shows up later or less often, and the mismatch produces no error, just a silently empty result for whatever it wasn't built to catch. A frequent failure mode in document and data extraction work (bank statements, invoices, structured PDFs, log parsing) where a field's format changes partway through a dataset, for example an account number that becomes partially masked after a certain date, and the extraction logic was written against the earlier format only. It looks exactly like a clean, successful run, because nothing errors and nothing gets flagged. The only way to catch it is to check the output itself for fields that are unexpectedly blank, rather than trusting a report that says zero issues were found.

Silent truncation

Silent truncation happens when an output ends early but still looks complete and reports success. A quick skim may miss it because the missing ending is not there to raise an alarm. Compare length or counts with the original and check the last meaningful piece of content.

AI, agents and automation

Skill

A reusable set of instructions an AI system loads when a matching situation appears. A skill defines how a particular kind of task should be done, encoding standards and procedure so the same work is performed consistently rather than improvised each time. Skills do not run on their own; they shape behavior when relevant work arrives.

Scheduled task

An instruction set that runs automatically on a clock, whether or not anyone is present. Scheduled tasks are where automation delivers compounding value and also where it fails most quietly, because nobody is watching at the moment they run. Any scheduled task worth having is worth monitoring for output rather than execution.

Cron job

A task set to run automatically at a fixed time or on a recurring schedule, defined by a compact five-part time pattern, minute, hour, day of month, month, day of week, known as a cron expression, for example one meaning "every day at 6:45 AM." The name comes from cron, the decades-old Unix scheduler this pattern originated in. It's now the standard way automated systems, including AI scheduled tasks, express recurring timing.

Agent

An AI system given a goal, a set of tools, and permission to decide its own steps, rather than following a fixed script. Agents are more capable than scripted automation and correspondingly harder to supervise, because the path they take is not known in advance. The practical requirement is not smarter agents but clearer boundaries: defined scope, defined tools, and a defined point where a human decides.

MCP (Model Context Protocol)

An open standard that lets AI systems connect to external tools and data sources such as email, calendars, file storage, and business applications. It matters because it turns an AI from something that talks about your work into something that can act on it. It also means access, permissions, and scope become real operational concerns rather than theoretical ones.

Connector

An authorized link between an AI system and an external account. Connectors are scoped and revocable, which is what makes them safe, and their scope is frequently narrower than people assume. A common source of confusion is an AI reporting that something does not exist when it simply falls outside what that particular connection was authorized to see.

Instruction layer

Where standing rules for an AI system live. Most platforms have several layers: global instructions that apply everywhere, project or workspace instructions with narrower scope, and reference material loaded only on demand. Knowing which layer is guaranteed to load is essential, because a rule written into a layer that does not load is documented rather than active, and it will appear to be working right up until it matters.

Context window

How much information an AI system can hold in working memory at once. When it fills, earlier detail degrades, which is why long sessions drift, repeat themselves, or lose established decisions. Practical implication: durable decisions belong in a document, not in a conversation.

Hallucination

Output stated confidently that is not grounded in any real source. The risk is not that AI systems are wrong sometimes, it is that wrong output arrives with the same fluency and confidence as correct output. This is why verification steps and source citation matter more than model quality in most business deployments.

Silent drift

An automated process still faithfully following instructions that no longer match the document meant to govern it. Nothing errors and nothing alerts. The procedure and the practice simply diverge over time, and the gap is usually discovered by accident, often long after it started causing damage. Preventing it requires periodically checking the live system against its documentation rather than assuming they match.

Parallel session

Two or more AI sessions working at the same time on related material. Sessions do not share state, so each can act on the same resource without knowing the other exists, including undoing each other's work. As organizations deploy more agents, this becomes a foundational design question rather than an edge case: what is the shared source of truth, and how does one agent learn what another already changed?

Human in the loop

A required human approval before an action that is difficult or impossible to reverse. The practical rule is to place approval gates around anything that sends, spends, publishes, or deletes, and to let everything else run unattended. The goal is not to supervise AI constantly, it is to be deliberate about which decisions stay human.

Grep/Glob

The two-tool pattern behind targeted retrieval. One tool finds files by name or pattern; the other searches inside file contents for a match. Together they let an AI system have broad access to a folder of information without pulling all of it into working memory at once. This is the mechanism that makes wide access and a manageable context window compatible: an AI can reach everything in a connected folder while only the specific files relevant to the question at hand are actually read.

Belt-and-suspenders (b-a-s)

Manually attaching a document to an AI project or workspace even though the AI already has broader access that would let it find that same document on its own. The redundancy is deliberate: it guarantees a specific piece of context loads automatically at the start of every session, rather than depending on the AI to retrieve it when relevant. Worth reserving for a small set of genuinely load-bearing documents, since doing it for everything defeats the purpose of having broad access at all.

Subagent

A specialized agent that a primary agent creates to handle one bounded piece of work, with its own working memory and a narrower set of tools, before reporting its result back and closing. Subagents exist mainly to protect the parent agent's own working memory: the exploratory searching, reading, and trial-and-error involved in a subtask stay contained inside the subagent rather than crowding out everything else the parent needs to remember. A useful mental model is a specialist a manager delegates a narrow task to, rather than doing the research personally.

Plugin

A distributable bundle that packages skills, specialized agents, and connections to outside systems together so they can be installed as a single unit. A plugin is the package; a skill is one instruction set inside it. Plugins are how AI capability spreads in practice: rather than building a capability from scratch, an organization installs a plugin someone else built and gets its skills, agents, and connections all at once.

GPT (OpenAI custom GPT)

OpenAI's version of a saved, custom-configured AI assistant: built from custom instructions, uploaded reference material, and a chosen set of tools, layered on top of ChatGPT. It is closer to a saved project setup than to a fully autonomous agent, since it is mostly a persona and a knowledge scope rather than something that takes multi-step action on its own. OpenAI has begun evolving this concept toward more autonomous agent products, which is the direction the whole industry is heading.

Gem (Google Gemini)

Google's version of a saved, custom-configured AI assistant inside Gemini, built from a name, instructions, and reference material. It serves the same basic purpose as OpenAI's custom GPTs: save a configuration once instead of re-explaining context every conversation. Of the major platforms' equivalents, it is generally the lightest-weight, closer to a saved prompt than to an autonomous agent that takes action on its own.

Seed message

A short, paste-ready block of the essential facts from a finished session or project, written so a person can drop it into a brand-new AI conversation and have the assistant pick up right where things left off, without re-explaining everything from scratch. The same idea shows up across the AI industry under different names: some call it a 'seed,' others describe it as part of 'context engineering' or 'context rehydration.' Whatever the name, the practice is the same: rather than carrying a full conversation history forward (which eventually overflows or degrades), carry forward only the load-bearing facts a fresh conversation actually needs.

AI-assisted software development

An AI assistant can draft code, suggest fixes, or explain unfamiliar parts of a system. The team still sets the requirements, reviews the result, tests the behavior, and decides what to release. The quality of that review matters more than whether a person or an AI typed the first draft.

C2PA / Content Credentials

An open standard for attaching tamper-evident provenance to a media file, recording where it came from and what has been done to it since. Distinct from an invisible watermark: the watermark lives inside the content itself, the credential lives alongside it as signed metadata. Both are becoming default on generated media, and both are worth understanding before assuming a generated asset can pass as a photograph.

Capture journal

A saved record of draft edits that can be recovered after a restart. It helps an automated system preserve work before submitting it for review. The record shows what was captured; it does not establish that the draft was approved or delivered.

Generally available vs preview

Whether a vendor considers a feature finished and supported, or still under test. The distinction is easy to skim past because both are available on the same account and often cost the same. Preview endpoints can change shape without notice, and a vendor's commercial protections, including copyright indemnity, usually do not extend to them. Anything load-bearing should be checked for which side of that line it sits on.

Generative media

AI that produces pictures, video, music or speech rather than text. It behaves differently to a chat model in three ways that matter commercially: it is priced per second, per image or per generation rather than per word, it often takes minutes rather than seconds to return, and its licensing terms are frequently stricter and less uniform than those covering text.

Groups (Claude)

Groups is an organizational feature on Claude's Enterprise plan that lets an administrator bundle members by team or department, then manage spend limits, permissions, and access to shared resources for the whole group at once rather than person by person. It sits above Projects, which organizes a body of work — Groups is the layer that organizes people and what they're allowed to touch. Worth knowing before rolling AI access out past a handful of users.

Idempotency

Being able to repeat an operation without doing its effect twice. This matters when an automated system loses the response and cannot tell whether an action succeeded. A reliable retry recognizes the earlier action instead of creating a duplicate.

Image-to-video

Giving a video generation model an existing still image to animate, rather than describing a scene from scratch and letting the model invent it. This matters more than it sounds in commercial work. Text-to-video produces something new every time, which means a client is approving a fresh subject with each attempt. Image-to-video anchors the motion to a frame already signed off, so what moves is the thing everyone agreed on.

OCR (optical character recognition)

Optical character recognition turns a picture of text, such as a scan, into characters a computer can search and process. It saves retyping, but lookalike characters and poor scans can produce errors. Verify names, amounts, and exact reference numbers against the original page.

Parse (document parsing)

Document parsing converts a file into organized text or data that software can work with. It is a translation of the original, so tables, labels, or whole passages can change or disappear. Compare important facts and completeness against the source before relying on the result.

Parsing tier

A parsing tier is a price and capability level offered by a document-processing service. Paying for a higher tier may improve one kind of extraction without improving every kind. Recheck both accuracy and completeness whenever you change methods or tiers.

Polling

Polling means asking a service at intervals whether a job has finished. It is useful when a task takes longer than one request. Set a reasonable interval and a deadline so a stuck job does not wait forever or send unnecessary requests.

Provider adapter

A design pattern for building on AI without being married to one supplier. Rather than writing code that calls a named vendor, you define the capability you want and put a thin translation layer behind it for each provider. Swapping or adding a supplier then means writing one small adapter instead of rewriting everything that depends on it. Given how fast models change and how often pricing and terms move with them, this is closer to basic hygiene than to over-engineering.

Revision ledger

A record that tracks the order of saved versions and which one has been confirmed at each step. It helps an automated process reject an older update that arrives late. A ledger is evidence of what a process recorded; it still needs reliable checks of the actual files and approval decisions.

SynthID

An invisible watermark embedded into AI-generated images, audio and video at the moment they are created. It is designed to survive cropping, compression and re-encoding, so the file stays identifiable as machine-generated long after it leaves the tool that made it. Worth knowing before generated material goes into anything public: the marker travels with the file, and treating generated work as indistinguishable from captured work is a bet against detection improving.

Token

A token is a small unit an AI system processes and often uses for usage limits or billing. It can be a word, part of a word, or another piece of input; the exact count depends on the model and format. Pictures and other media may also add to usage, so the same information can cost different amounts to process in different forms.

Vibe coding

Vibe coding usually means describing a program to AI, trying what it produces, and continuing by feel without closely reading the code. It can be a fast way to explore an idea. For software people rely on, reviewing the code, testing its behavior, and controlling releases remain separate responsibilities.

Client delivery and privacy

Allowlist

An allowlist starts with a closed door and names the items that may pass through it. For a document transfer, that might mean naming each approved file rather than copying an entire folder. The list still needs a trusted owner and a check that the transfer follows it.

Model output indemnity

A vendor's contractual promise to defend a customer against third-party copyright claims arising from output their model generated. The detail that catches people out is that it is granted model by model rather than vendor by vendor, and usually only for generally available versions rather than preview ones. Two models from the same company, billed to the same account, can sit on opposite sides of that line. Anyone putting generated material into commercial work should check which specific models are covered rather than assuming the vendor covers all of them.

Pre-download isolation

If a cloud service should see only selected documents, make a separate collection containing only those approved documents before connecting it. Removing private documents after the service downloads a larger collection does not undo that exposure.

Scratch asset

Material made to help a decision rather than to be delivered. A rough music bed used to test whether a film wants something warm or something spare is a scratch asset; the licensed track that eventually sits under it is a delivery asset. The distinction is worth making explicit and deciding once, because generated material is convincing enough that the line blurs under deadline, and the cost of blurring it falls on whoever handed the work over.

Sensitivity classification

Sensitivity classification answers a practical question: who may see this exact information? A folder name can be a clue, but individual documents may need different treatment. The decision should come from an authorized reviewer and be checked before sharing.

Identity, credentials and access

API key

An API key is a secret credential that lets software use a service. Treat it like a password: anyone who obtains it may be able to act with its permissions. Keep it out of shared documents and chat, limit what it can access, and update the programs that use it when the key changes.

Git and version control

Merge

Adding a proposed set of changes into the shared version of a project. Teams commonly review a pull request before merging it. A merge changes the project record; delivering that change to a running app or another device is a separate step.

Pull request (PR)

A proposal to add changes to a project. It shows what changed, keeps review discussion together, and can run automated checks. Accepting the proposal means merging it into the shared version; opening a pull request alone does not change that version.

Definitions on this page come from BK Blueprint's working glossary, with separate public wording reviewed before publication. Structured term data helps search and AI systems read each definition. Website updates are checked against the accepted source.

Every definition, one file.
Download as Markdown