Building BHM™: Why Observation Wasn't Enough
Resource: The Blackwell-Hart Methodology™ (BHM™)
Series: Building BHM™
A methodology cannot be built from observation alone.
Observation is where the process begins. It allows a condition to be seen, recorded, and described. But observation by itself does not establish whether what was observed is meaningful, repeatable, persistent, or connected to a particular intervention. Once BHM™ began being applied to real-world AI-assisted discovery systems, that distinction became increasingly important.
AI-generated outputs are not static. They can vary according to the model being used, the structure of the question, the information available to the system, retrieval conditions, geographic context, system updates, and other factors that may sit outside the control of the person conducting the observation. A single response can therefore tell us what happened under a particular set of conditions, but it cannot, by itself, tell us why it happened or whether the same result would occur again.
That created the next methodological problem.
From Observation to Evidence
If an observation is going to contribute to a methodology, it must be possible to establish what actually happened and the conditions under which it happened.
This means recording the AI output itself, the prompt that produced it, the entity being tested, the model used, the date and time of the observation, and the relevant outcome. Depending on the test, that outcome might include whether an entity was included, how it was categorized, whether it appeared in a recommendation, whether a previously recurring reference disappeared, or whether the system described the entity differently from an earlier observation.
The evidence does not necessarily come from one source. A screenshot establishes one observation. A timestamp establishes when that observation occurred. A repeated test provides another observation under the same defined conditions. A second model provides a comparison point. A later measurement period provides evidence about persistence.
Each piece of evidence adds context to the original observation.
Observation records an outcome. Evidence establishes the conditions under which that outcome was observed.
The Problem with the Single Result
A single unexpected response can be important. It can reveal something that warrants investigation. It cannot automatically establish causation.
Suppose an entity appears in an AI-generated recommendation after a website has been changed. The timing may be interesting, but several explanations remain possible. The result may be associated with the changed information. It may instead reflect different query wording, a change in retrieval conditions, a model update, an external reference changing, an alteration to the entity itself, or ordinary variation between runs.
The same problem exists when an apparently positive result disappears.
Without a defined observation process, there is no reliable way to distinguish a meaningful change from an isolated result.
BHM™ therefore had to move beyond asking, What happened?
It had to begin asking, Under what conditions did it happen, and can the observation be examined again?
The Need for a Controlled Condition
BHM™ does not control AI systems internally. It does not control model weights, retrieval mechanisms, ranking processes, training data, or system updates.
What can be controlled is the observation environment.
A structured test can specify the entity being examined, the prompt being used, the model being tested, the number of repeated runs, the date of the observation, and the categories of outcome being recorded. This does not turn an AI-assisted discovery system into a controlled laboratory environment. It does, however, create a defined basis for comparison.
That distinction matters.
The objective is not to eliminate every source of variation. It is to make the conditions surrounding an observation sufficiently clear that differences between observations can be examined rather than simply assumed to be meaningful.
From “What Happened?” to “What Changed?”
As BHM™ developed, another distinction became necessary.
Not every observable change represented the same kind of change.
An entity could become more discoverable without becoming more consistently recognized. It could be recognized correctly without being selected as a recommendation. It could be recommended without the underlying authority infrastructure becoming stable. A citation could appear without demonstrating that the system had adopted the entity's own definition or conceptual structure.
Visibility was therefore not one thing.
BHM™ began separating observable conditions such as recognition, recommendation, authority, and source-of-truth alignment rather than treating them as interchangeable indicators of success.
This separation made measurement possible.
The Development of Measurement
Once the observations were being recorded consistently, BHM™ needed to define what would actually count as a measurable change.
For example, a test could establish criteria for whether an entity was included, whether it was associated with the intended category, whether it appeared in a leading recommendation position, whether relevant citations recurred across observations, or whether the system's description remained consistent over time.
The purpose of introducing measurement was not to turn AI interpretation into a collection of arbitrary numbers.
The purpose was comparability.
If the same condition can be observed repeatedly, and the same type of outcome can be recorded each time, then observations can begin to be compared across prompts, models, runs, and measurement periods.
That creates a very different evidentiary position from simply saying that an AI system “seemed to understand” an entity better.
Measurement Does Not Create Certainty
Measurement does not reveal what is happening inside an AI model.
It does not expose hidden reasoning, establish the internal cause of a response, or remove the uncertainty inherent in observing dynamic systems.
What measurement does provide is a more disciplined way to examine what can actually be observed.
A result that varies is still a result. A result that fails to repeat is still information. A result that appears consistently across defined conditions provides a different level of evidence from an isolated response.
Variation is therefore not something that BHM™ needs to hide.
It is something that needs to be recorded.
The Role of the Baseline
This also established the importance of a baseline.
A BHM™ baseline is the observed starting condition against which later observations can be compared. It is not a universal industry benchmark and does not represent an assumed level of performance that every organization should achieve.
The baseline belongs to the entity being examined.
If an entity is absent from a defined set of non-branded prompts before implementation and appears repeatedly afterward, that change can be documented. If it is already consistently recognized before implementation, a different question may need to be measured.
Without the baseline, however, later observations lose an important reference point.
The methodology therefore needed to establish the condition before asking whether the condition had changed.
Why Repeated Observation Matters
Repeated observation provides another layer of evidence.
If a result appears once, it may represent an isolated outcome. If the same result appears repeatedly under the same defined conditions, it becomes more reasonable to treat it as a recurring pattern. If the pattern also appears across more than one AI model or persists across separate measurement periods, the evidentiary position becomes stronger again.
None of this makes the underlying system deterministic.
It simply reduces the likelihood that the methodology is relying on a single observation.
That distinction is central to BHM™.
Repeated results do not prove why a system produced a particular response. They provide stronger evidence that the observed condition was not limited to one isolated run.
From Evidence to Measurement
The methodological progression was therefore becoming clear:
Observation → Evidence → Measurement
Observation identifies what occurred.
Evidence establishes the circumstances in which it occurred.
Measurement makes comparable change possible.
Numbers became useful only after those foundations were established. A percentage, frequency, position, or recurrence count can describe an observed condition, but the number itself is not the methodology.
Numbers are evidence. They are not, by themselves, methodology.
What BHM™ Had to Measure
The development of BHM™ required measurement across several distinct dimensions rather than a single generalized concept of “AI visibility.”
These dimensions developed into:
Discoverability — whether the entity could be retrieved or included under defined conditions.
Category Recognition — whether the entity was consistently associated with the intended category or role.
Recommendation Positioning — whether the entity appeared in leading recommendation positions under defined conditions.
Infrastructure Stability — whether relevant recognition, citation, routing, and categorization patterns persisted over time.
Reasoning Alignment — whether generated descriptions increasingly aligned with the definitions, concepts, frameworks, and reference materials associated with the entity.
These dimensions were not intended to expose internal model behavior. They were designed to describe observable outcomes.
That distinction remains fundamental.
When a Methodology Becomes Examinable
By this point, BHM™ was no longer simply documenting interesting AI responses.
It had defined conditions, observations, evidence, measurements, and limitations.
That changed the nature of the work.
An AI-generated output that had not been independently repeated remained an observation or hypothesis. A result that could be reproduced under defined conditions provided stronger evidence. A result that persisted across models and measurement periods provided another level of evidentiary support.
The methodology was becoming something that another person could examine.
That was necessary because a methodology cannot depend entirely upon the authority of the person who created it. Its observations, definitions, measurements, and limitations need to be visible enough that the process itself can be questioned.
The Next Problem: Evaluation
Measurement solved one problem, but it created another.
Once BHM™ could describe that a condition had changed, the next question was what that change actually meant.
An increase in inclusion is a measurable observation. Whether that increase represents meaningful progress depends on the conditions being evaluated, the baseline being used, the consistency of the result, and the purpose of the test.
Measurement could describe movement.
It could not, by itself, determine significance.
That required the next methodological layer: evaluation.
And that is where BHM™ continued to evolve.