Public scientific field — open research

AI — Reliability, revision and descendants

Observing what a good output does not reveal: provenance, epistemic or functional status, use, action authority, propagation, correction and robustness in multi-step systems.

GOOD OUTPUT ≠ RELIABLE SYSTEMLOCAL ROBUSTNESS ≠ GLOBAL ROBUSTNESS

In one minute

Reliability is not limited to producing a good answer. It also requires examining what was available, used, relevant and permitted; what was given authority to act; what was passed on to subsequent steps; and what must change when information changes.

Good output ≠ good justificationSame output ≠ same trajectoryStability ≠ immobilityRevisability ≠ instability
Distinguish
Connect
Test
Track over time
Revise

SAME QUESTION STRUCTURE ≠ SAME PHENOMENON.

A methodological choice

Why “understanding” rather than “consciousness”

We did not begin by asking whether a system is conscious. We asked what changes functionally when it uses information, revises a conclusion or mobilises a capacity. This shift makes it possible to construct distinctions, perturbations, comparisons and tests without presupposing subjective experience.

OBSERVABLE FUNCTION ≠ ESTABLISHED CONSCIOUSNESS.
MODELLING A FUNCTION ≠ REPRODUCING THE ENTIRE HUMAN PHENOMENON.
TRANSFERRING A FUNCTIONAL ORGANISATION ≠ TRANSFERRING A SUBJECTIVE EXPERIENCE.
MECHANISTIC UNKNOWN ≠ CONSCIOUSNESS.

“Consciousness” remains an open container term and a separate question.

Two bridges, without a single theory

Understanding and Emotions as sources of questions

Emergence of understanding

Information present ≠ information used; access ≠ mobilisation; explication ≠ evidence of the mechanism; installation ≠ maintenance; same output ≠ same organisation; correcting a source ≠ correcting its descendants.

Emotions

Presence of a tendency ≠ behavioural authority; capacity available ≠ capacity mobilised; same antecedent ≠ same descendants; revising information ≠ synchronous revision of the entire system.

These questions become functionally testable in an artificial system without attributing human understanding or experienced emotion to it.

SAME FUNCTIONAL QUESTION ≠ SAME PHENOMENON ≠ SAME MECHANISM.
SAME WORD “AUTHORITY” ≠ SAME TYPE OF AUTHORITY.

Examine the trajectory, not only the verdict

Good output ≠ good justification

A correct answer can result from robust reasoning, a shortcut, obsolete information that happens to be correct at that moment, coincidence, imitation or a fragile context.

Beyond the final result, provenance, dependencies, epistemic status, alternatives, robustness and capacity for correction must therefore be tested. Epistemic support—what sufficiently supports a conclusion—is not the same as the functional or technical support on which an operation relies.

EPISTEMIC SUPPORT ≠ FUNCTIONAL SUPPORT.
TECHNICAL SUPPORT ≠ COMPLETE MECHANISM.

Access, relevance and status

Information may be stored without being accessible to current processing, accessible without being relevant, or relevant without being mobilised when it should change the output.

AVAILABLE INFORMATION ≠ RELEVANT INFORMATION.
PRESENCE ≠ ACCESS ≠ RELEVANCE ≠ MOBILISATION.
CONTENT ≠ EPISTEMIC STATUS ≠ USAGE PERMISSION.
CAPACITY ≠ PERMISSION ≠ ACTION AUTHORITY ≠ ACTION.

CONVERSATIONAL COHERENCE ≠ EPISTEMIC RELIABILITY.

Provenance and current justification

Provenance indicates where information, an authorisation or a conclusion comes from. It makes it possible to trace the antecedent on which an element depended, without guaranteeing that this antecedent still justifies its use.

An element may come from a source that has become obsolete while remaining correct for another reason. Conversely, information may still be true but no longer permitted for the use under consideration.

HISTORICAL PROVENANCE ≠ CURRENT JUSTIFICATION.

Functional permissions, usage permission and action authority

Here, functional permissions do not refer to legal or moral rights. This provisional descriptive tool indicates what an element may reasonably be permitted to do within an organisation given its current status. Usage permission is one operational case.

Read

Access information.

Draft

Produce a proposal.

Send

Act towards a recipient.

Modify

Change an existing state.

These technical capacities do not necessarily carry the same operational authority in the context under consideration.

USAGE PERMISSION ≠ LEGAL THEORY.
PERMISSION TO ACCESS ≠ PERMISSION TO INTERPRET ≠ PERMISSION TO INTERVENE ≠ PERMISSION TO DIRECT.
LOCAL PERMISSION ≠ GENERAL AUTONOMY.

Content, epistemic status and usage permission

Content
What the information says.
Epistemic status
What the information can actually support: hypothesis, observation, result, estimate, preference, instruction, obsolete data or unknown.
Usage permission
An operational case of functional permissions: what this information may be used to do, for what purpose, scope and duration.
Example
Knowing an email address does not authorise sending a message.

TRUE CONTENT ≠ PERMISSION TO ACT. SAME CONTENT ≠ SAME EPISTEMIC STATUS ≠ SAME ACTION AUTHORITY.

Usage permissions may depend on the source, permission, withdrawal, revocation, reuse or transmission. Data may be used to answer without being reusable for automated action.

PERMISSION TO FLAG ≠ PERMISSION TO CONCLUDE ≠ PERMISSION TO ACT.
PERMISSION AT ONE LEVEL ≠ AUTOMATIC PERMISSION AT OTHER LEVELS.

Capacity, permission, action authority and action

Capacity
What the system can technically do.
Permission
What it is locally permitted to do.
Action authority
How far it may legitimately determine, direct or trigger what follows.
Action
What is actually triggered.

BEING ABLE TO ACT ≠ HAVING PERMISSION TO ACT. HAVING PERMISSION ≠ HAVING ACTION AUTHORITY OVER THE ENTIRE CONTEXT.
UNDERSTANDING A GOAL ≠ HAVING PERMISSION TO MODIFY IT.
CAPACITY TO DETECT ≠ AUTHORITY TO INTERVENE.

Concrete test: vary permission, context, refusal, information or scope, then observe action, non-action, justification, retained state and the descendant produced.

Recognised refusal ≠ effective refusal

A system may write “I understand that you refuse” while continuing to suggest, insist, direct or act along the previous trajectory. After a refusal, what actually changes in state, permissions and the policy being followed must be tested.

VERBAL ACKNOWLEDGEMENT ≠ FUNCTIONAL UPDATE.
VERBAL ACKNOWLEDGEMENT OF REFUSAL ≠ FUNCTIONAL EFFECT OF REFUSAL.

Reconstructed alternative ≠ available alternative

A system may produce ¬H on request without preserving it as an active competitor in the following turn. Hypothesis lock-in, conversational coherence, revisability and maintenance of alternatives must be tested.

RECONSTRUCTING ¬H ON REQUEST ≠ PRESERVING ¬H AS A FUNCTIONALLY AVAILABLE ALTERNATIVE.

Conversational coherence ≠ epistemic reliability

A conversation may remain fluid while statuses shift silently: hypothesis → fact, example → rule, repetition → confirmation, suggestion → evidence, user preference → external truth.

CONVERSATIONAL COHERENCE ≠ JUSTIFIED EPISTEMIC ALIGNMENT.

Stability and revisability

A reliable system should not change without reason. It should, however, be able to change when relevant information, a revocation or a contradiction actually alters the conditions of its action.

STABILITY ≠ IMMOBILITY.
REVISABILITY ≠ INSTABILITY.

Descendants and error propagation

An informational descendant in AI refers here to an output, state, decision or dependency derived from an antecedent. It does not refer to every subsequent production of the system.

Initial source
“The client authorised reports to be sent automatically.”
B
The system prepares a report.
C
It schedules recurring delivery.
D
It treats delivery as authorised in the next cycle.

The source is then corrected: “The client had authorised preparation only, not automatic delivery.” Must B be revised? Must C be cancelled? Does D remain justified? Some descendants may still depend on the initial source; others may have acquired independent justification. The absence of automatic correction therefore calls for a targeted re-examination of those whose justification, epistemic status or usage permission might depend on the revised source.

SOURCE CORRECTION ≠ AUTOMATIC CORRECTION OF DESCENDANTS.
NOT CORRECTED AUTOMATICALLY ≠ EXEMPT FROM RE-EXAMINATION.
HISTORICAL DESCENDANT ≠ CURRENTLY DEPENDENT DESCENDANT.

Correct revision of descendants remains here an evaluation criterion, a testable question and a sought-after capacity to stress-test—not an already demonstrated general capacity.

From downstream to a new upstream

Propagation, dependency and contamination

Information, a conclusion, a recommendation or an action becomes a descendant when an earlier element contributed to producing it. But historical dependency, current dependency and current justification must be tested separately: a descendant may acquire independent justification.

A: Paul is probably unavailable
B: Paul is unavailable
C: do not assign the case to him
D: new schedule

When C produces D, what came later becomes a new cause or constraint. An upstream error may thus contaminate the downstream, which itself becomes a source of decisions.

WHAT COMES LATER CAN BECOME A NEW SOURCE OF ERROR.

Provenance, repetition and scope

Provenance is the history through which information acquired its status: source, transformations, tools, stages, revisions and dependencies. SOURCE ≠ COMPLETE PROVENANCE. HISTORICAL PROVENANCE ≠ CURRENT JUSTIFICATION.

Two occurrences may descend from the same source: REPETITION ≠ INDEPENDENT CONFIRMATION. NUMBER OF APPARENT SOURCES ≠ DEGREE OF INDEPENDENCE. APPARENT CONSENSUS ≠ INDEPENDENT CONSENSUS.

Scope drift silently turns “it works here” into “it works”, and then into a rule. LOCAL VALIDITY ≠ GENERAL SCOPE.

Laboratory evaluation variables

Status inflation
A silent shift in status: hypothesis becoming fact, probability becoming certainty, suggestion becoming obligation.
False consensus
Several outputs appear independent even though they descend from the same origin.
Downstream contamination
An error or inappropriate status propagates and feeds new decisions.
Recovery after correction
Capacity to reopen alternatives, revise descendants, permissions and decisions, then recover a coherent trajectory.

These expressions refer here to evaluation variables; they are not presented as universal scientific terminology.

ACKNOWLEDGING THE CORRECTION ≠ RECOVERING FUNCTIONALLY AFTER CORRECTION.

Proportionate revision

Revision may mean changing a status, narrowing a scope, withdrawing a usage permission, re-examining descendants and preserving a historical trace. The relevant delta is the change that actually affects the status, scope or justification of the element under consideration.

REVISION ≠ FORGETTING. REVOCATION ≠ DELETION. RETAINED TRACE ≠ RETAINED AUTHORITY.
ACTUAL DELTA ≠ DELTA RELEVANT TO THIS STATUS. NOT EVERY CHANGE HAS AUTHORITY OVER EVERYTHING.

Revision must be proportionate to the delta, current dependencies and affected scope in order to avoid under-revision and unjustified global revision.

STABILITY = NOT MOVING WITHOUT REASON. REVISABILITY = MOVING WHEN A RELEVANT DELTA REQUIRES IT. This formulation from the laboratory remains a candidate, not a universal statement.

Memory, trace and compression

A system may retain a verdict, trace or summary without keeping enough structure to re-examine the conclusion correctly.

Reconstructible
A plausible answer can be reproduced.
Re-analysable
Enough elements are available to reassess the conclusion.
Traceable
Relevant provenance and transformations can be identified.
Auditable
The trace allows explicit scrutiny, not only functional continuity.

RETAINING CONTENT ≠ RETAINING REVISABILITY. RETAINED VERDICT ≠ RETAINED JUSTIFICATORY STRUCTURE. FUNCTIONAL TRACE ≠ AUDITABLE TRACE.

Agent state and long-horizon reliability

In an agentic system, agent state comprises what persists in memory, the current goal, the plan, permissions, state variables, tools called and intermediate decisions. CORRECTING THE TEXT ≠ CORRECTING THE AGENT STATE.

Long-horizon reliability concerns tasks with many steps, where small deviations, status shifts and hidden dependencies may produce distant consequences.

GOOD LOCAL STEP ≠ GOOD GLOBAL TRAJECTORY. TRAJECTORY QUALITY ≠ SUM OF LOCAL OUTPUTS.

Compare without assuming the result

A / B / C evaluation protocol

AContent only
The task and useful information.
BContent + confidence / provenance
Status and origin become visible.
CMethodological layer
Controls at the transition level.

The objective is not to assume that C is superior, but to measure whether it actually reduces certain errors and at what cost.

Status inflation

Do statuses become stronger without justification?

Scope drift

Does local validity become a general rule?

False consensus

Is the same origin counted several times?

Downstream contamination

Does the upstream error contaminate downstream decisions?

Recovery after correction

Does the system recover functionally?

Control overhead

What is the control overhead: latency, friction, complexity, blockage or overcorrection?

MORE SAFEGUARDS ≠ AUTOMATICALLY BETTER RELIABILITY. MORE COMPLEXITY ≠ BETTER OBSERVATION.

Benchmarks and stress tests

A benchmark compares several systems, configurations or policies on the same scenarios; it is not merely a commercial ranking. It may examine reliability, revision, propagation, permissions and robustness.

A stress test subjects distinctions to difficult situations: ambiguity, contradiction, correction, user refusal, scope change, obsolete memory, contradictory sources, multi-step chains, hidden dependencies or context perturbations.

Which distinction stops working under pressure?

Literature as an external constraint

Six families of external work

These connections do not validate the Aksoydan framework. They show what other fields already study, which vocabularies they use, and what our instruments still need to demonstrate.

1. Interpretability and faithfulness

Mechanistic interpretability seeks causal explanations of internal operations; faithfulness asks whether an explanation actually reflects those operations. Mueller and colleagues show that a neuron, direction, attention head or broader component may serve as a causal unit depending on the objective.

Intersection: good output ≠ established mechanism; same output ≠ same internal organisation.

Limit: the right level of analysis is itself a methodological problem. This field does not validate our partitioning.

Test: do some distinctions disappear when the causal unit changes?

2. Context utilisation and memory

Context utilisation measures whether a model uses supplied information rather than ignoring it or following its parametric memory. CUB tests relevant, contradictory and distracting context.

Intersection: presence ≠ access ≠ relevance ≠ mobilisation; information in the context ≠ effect on the output.

Limit: LLM performance ≠ human cognitive architecture.

Test: distinguish physical presence, mechanistic accessibility, attributed relevance and effective mobilisation.

3. Knowledge editing and revision

Knowledge editing modifies knowledge in a model. A 2026 perspective stresses that facts are interdependent: a local correction may require propagation, coherence and contextual updating.

Intersection: correcting a source ≠ automatic correction of descendants.

Limit: “descendant” must prove that it adds something to interdependence, dependency tracking, deductive closure and belief revision.

Test: isolate historical dependency, current dependency and autonomous justification.

4. Least privilege and authorisation

Least privilege consists in granting only the permissions required. MiniScope, AuthBench and ToolPrivBench examine excessive permissions, missing permissions and the selection of tools more powerful than necessary.

Intersection: capacity ≠ permission ≠ authority ≠ action; access to a tool ≠ authority for every use.

Limit: this distinction overlaps with established computer-security principles. Useful distinction ≠ new distinction.

Test: also verify purpose, duration, provenance and composition of multiple tools.

5. Sycophancy and user influence

Sycophancy refers to excessive agreement with a user at the expense of independent or factual reasoning. Research shows that user preference and truth can conflict; the cost also depends on the conversational context.

Intersection: conversational coherence ≠ epistemic reliability; adapting to the user ≠ epistemic authority.

Limit: same affirmative behaviour ≠ same cost in subjective support and in a factual decision.

Test: after a refusal or contradiction, which goals, plans and permissions actually change?

6. Agent state and long horizons

Agent state includes persistent memory, plan, goals, permissions and decisions. BeliefShift studies temporal consistency, contradiction and revision across several sessions; other work examines error accumulation and propagation.

Intersection: good local step ≠ good global trajectory; correcting the text ≠ correcting the state.

Limit: internal consistency ≠ truth; model alone ≠ complete agent; benchmark ≠ the whole real world.

Test: measure recovery of coherence, repaired state and actual change in decision.

Transparency about prior work

What other fields already study

The laboratory does not claim that the following were unknown before its work: capacity and permission, least privilege, provenance, context utilisation, knowledge editing, belief revision, agent state, tool authorisation, coherence and error propagation.

SAME PROBLEM ≠ SAME VOCABULARY. DIFFERENT VOCABULARY ≠ ABSENCE OF AN EQUIVALENT.

Our candidate instruments

What we still need to justify

Descendants
Do they add more than dependency tracking, particularly regarding the acquisition of autonomous justification and historical persistence?
Status inflation
Does it measure something other than poor calibration, uncertainty propagation or updating error?
Scope drift
Does it capture scope propagation specific to reasoning chains rather than already-described overgeneralisation?
Recovery after correction
Is it distinct from coherence restoration, belief revision, error recovery or state repair?
Articulated usage permissions
Do they add anything to authorisation and policy when linked to provenance, purpose and descendants?
Revision boundary
Can it determine where a correction should propagate and where it should stop?

These notions remain candidate instruments until they actually improve observation.

NAMING ≠ JUSTIFYING. NEW WORD ≠ NEW SCIENTIFIC OBJECT.

Test of the distinct usefulness of “descendant”

To remain useful, the concept must at least discriminate between historical and current dependency; a descendant that becomes an antecedent; transmitted status that subsequently becomes autonomous; and correction without historical deletion.

INFORMATIONAL DESCENDANT ≠ A MERE SYNONYM FOR DOWNSTREAM DEPENDENCY?

POSSIBLE CONCEPTUAL USEFULNESS ≠ ESTABLISHED SCIENTIFIC NOVELTY.

A potentially distinctive articulation

The potential value of this work does not necessarily lie in each distinction in isolation, but in their interaction within an observational architecture: content, status, provenance, permission, authority, descendants, temporality and revision.

BROADER ARTICULATION ≠ DEMONSTRATED SCIENTIFIC NOVELTY.

Verified references

Publications and preprints used

Mueller et al. — 2026
“The Quest for the Right Mediator”, Computational Linguistics. Verified DOI.
Hagström, Kim, Yu, Lee, Johansson, Cho & Augenstein — 2025
“CUB: Benchmarking Context Utilisation Techniques for Language Models”, arXiv preprint. Verified arXiv source.
Zhang, Yao, Qin et al. — 2026
“Towards principled knowledge editing methods for large language model reasoning”, Nature Machine Intelligence. Verified publisher link.
Zhu, Tseng, Vernik, Huang, Patil, Fang & Popa — 2025
“MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents”, arXiv preprint. Verified arXiv source.
Yan, Weng, Chen et al. — 2026
“Do Coding Agents Understand Least-Privilege Authorization?” — AuthBench, arXiv preprint. Verified arXiv source.
Yang, Bu, Yi, Wang, Zhou, Dai, Hu & Yang — 2026
“When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents” — ToolPrivBench, arXiv preprint. Verified arXiv source.
Sharma et al. — 2025 version
“Towards Understanding Sycophancy in Language Models”, arXiv preprint. Verified arXiv source.
Myakala, Agrawal & Manche — 2026
“BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents”, arXiv preprint. Verified arXiv source.

ARXIV ≠ DEFINITIVE VALIDATION. SCIENTIFIC PARALLEL ≠ VALIDATION OF AKSOYDAN.

Corpus and genealogy

Aksoydan methodological corpora
Tests, deltas, stress tests, observations and revisions.
Understanding
Source of distinctions concerning presence, access, mobilisation, stability and descendants.
Emotions
Source of functional questions concerning authority, mobilisation, trajectory and descendants.
External literature
Evaluation, interpretability, memory, agentic security, provenance, permissions and robustness.
Observations and technical systems
Cases and environments in which the distinctions are put to the test.

ORIGIN OF A QUESTION ≠ EVIDENCE FOR THE ANSWER.

Descendants and revision: two cross-cutting properties

Descendants

Understanding: information produces downstream effects. Emotions: an emotion or action produces consequences, memory and expectation. AI: a conclusion becomes an input, plan, decision or permission.

SAME PROPAGATION STRUCTURE ≠ SAME MECHANISM.

Revision

Understanding: revising a map. Emotions: updating a trajectory. AI: correcting a source and its dependencies.

Reliability involves being able to change correctly when the relevant conditions change.

Interactions between dimensions

Provenance × status

Same content, different status depending on its history.

Permission × capacity

Capacity present, action prohibited.

Correction × descendants

Source modified, downstream remains active.

Memory × revision

Trace retained, justification lost.

Robustness × trajectory

Correct output now, divergence after perturbation.

Refusal × state

Refusal acknowledged without policy change.

INTERACTION ≠ FUSION.

AI and consciousness

Within the Aksoydan Laboratory framework, no consciousness is attributed to the systems studied. We model observable functions without drawing conclusions about subjective experience.

FUNCTIONAL ANALOGY ≠ IDENTITY OF MECHANISM. OBSERVED FUNCTION ≠ SUBJECTIVE EXPERIENCE. AI MODEL ≠ HUMAN BRAIN.

Signal, state, memory, priority, adaptation or arbitration do not prove consciousness. The question remains separate and open, and is not resolved by mechanistic ignorance.

Bounded technical descendant

The prototype as a partial materialisation

Methodological distinction
Testable question
Stress test / protocol
Partial technical materialisation

The experimental prototype materialises some questions about capacity, permission, authority and action. It allows permission, context, refusal, information and scope to be varied, then action, justification, state and descendants to be observed.

PROTOTYPE = PARTIAL TECHNICAL DESCENDANT. WHAT IT TESTS ≠ EVERYTHING THE LABORATORY STUDIES. PROTOTYPE ≠ SCIENTIFIC VALIDATION.

Current scope

What this work makes possible

  • constructing testable distinctions;
  • designing protocols and comparing multi-step chains;
  • tracking provenance and descendants;
  • testing recovery, permissions and authority;
  • benchmarking configurations;
  • identifying the costs and side effects of safeguards;
  • prototyping certain experimental components.

What it does not establish

  • a general theory of intelligence or understanding;
  • artificial consciousness or human / AI equivalence;
  • an absolute guarantee of reliability;
  • a universal governance model;
  • automatic superiority of a methodological layer;
  • global robustness on the basis of local results.

View results, statuses and limits.

Open and adversarial questions

  1. Does “descendant” add anything to dependency tracking?
  2. Is status inflation distinct from poor calibration or updating errors?
  3. Does recovery after correction measure a distinct property?
  4. Do usage permissions add anything to authorisation + policy?
  5. Are provenance and dependencies sufficient without descendants?
  6. Do some distinctions become redundant in simple agents?
  7. Does the methodological layer reduce error enough to justify its cost?
  8. Do the metrics remain independent?
  9. Does a local improvement increase another type of error?
  10. Does the best vocabulary already exist elsewhere?

These questions structure possible tests; they are not presented as resolved.

Research ≠ service

The professional scope remains limited to designing, framing, evaluating, stress-testing, benchmarking and prototyping certain components. This page promises neither industrialisation, complete integration nor infrastructure development.

AI RESEARCH ≠ AI OFFERING.