# Asset Discovery Without Flattening History

**Public-candidate methodology note — review required**  
**Read-only discovery evidence, not an ownership or publication claim**

## The archive problem

A working archive rarely behaves like a clean portfolio library.

The reviewed corpus contained framework records, application code, governed trials, shipment mirrors, fieldbooks, visual experiments, music sessions, exports, and historical copies. The same file could appear in several delivery packages. A polished image could be an experiment. A filename could say `master`, `official`, or `v2` without any decision record establishing that status.

The task was not to gather everything attractive into one folder. It was to discover professional evidence without destroying the distinctions that made the evidence trustworthy.

## Scale and bounded scope

The read-only source census contained 8,474 pre-existing files, excluding `.DS_Store`:

| Area | Files |
|---|---:|
| AlexOS | 1,429 |
| Baron Creative | 59 |
| Baron Engineering | 6,891 |
| Model Testing | 5 |
| General Photos | 90 |
| **Total** | **8,474** |

The full census was not treated as a mandate to open every file.

- 3,353 entries matched the bounded artifact-oriented extension set after exclusions.
- 983 non-excluded document, image, code, and archive files received SHA-256 hashes.
- 32 non-journal-named images in the focused professional folder were visually reviewed.
- 30 path-matched sensitive items were excluded from content and hash review.
- audio was inventoried structurally but not played;
- personal photographs were counted but not visually inspected; and
- journals, chat archives, legal, tax, financial, credential, and secret-bearing sources remained protected.

This was asset discovery, not a complete semantic reading of 8,474 files.

## The method

### 1. Establish the found state

Before selecting candidates, record what the corpus already is: its top-level areas, file types, storage patterns, apparent project families, and protected zones.

This prevents the review from silently redesigning the archive while supposedly describing it. No source was reorganized, renamed, converted, deleted, published, deployed, or promoted.

### 2. Separate observation from interpretation

The review used four evidence labels:

- **Observation:** directly visible or structurally present.
- **Metadata evidence:** filesystem facts, embedded metadata, dimensions, archive members, or exact hashes.
- **Interpretation:** a bounded conclusion that may still be wrong.
- **Unresolved:** provenance or authority not supported by current evidence.

For example, a later modification date and larger file size can support “likely later iteration.” They cannot establish “approved design.”

### 3. Hash bounded, non-sensitive candidates

SHA-256 was used to identify exact byte-for-byte copies among 983 bounded files. The resulting map contained 79 exact-duplicate groups:

| Copies in group | Number of groups |
|---:|---:|
| 2 | 13 |
| 3 | 1 |
| 4 | 26 |
| 5 | 38 |
| 6 | 1 |

The dominant pattern was governed meeting and trial material mirrored across active folders, human-readable libraries, and dated shipment packages.

That is a copy relationship—not a supersession relationship.

### 4. Keep duplicates separate from versions

An exact hash can establish identical bytes. It cannot establish which copy is authoritative.

Different hashes are even less decisive:

- office containers may differ because of internal metadata;
- visually identical assets at different resolutions will not share a hash;
- a `v2` filename suggests sequence, not ratification; and
- multiple exports may descend from one design without preserving edit history.

The review therefore kept three questions separate:

1. Are these exact copies?
2. Are these probably related versions?
3. Is any one of them approved or current?

Only the first can be answered by a cryptographic hash alone.

### 5. Route by provenance and sensitivity

Candidate usefulness did not override privacy.

Professional documents, controlled creative experiments, personal sources, journal material, client-adjacent work, and secret-bearing records received different default lanes. When prompts, creators, edit histories, permissions, or outcomes could not be established, the record remained unresolved.

No identity inference was attempted for people in images. Likely generated imagery was not declared AI-generated without supporting prompt, tool, manifest, or participant evidence.

### 6. Select candidates without promoting them

A candidate register answered: “What deserves focused review next?”

It did not answer:

- what is true;
- what is canonical;
- what is public;
- what a client approved;
- what outcome occurred; or
- what should represent the final brand.

The package explicitly assigned itself no authority effect. Discovery created a review queue, not a release decision.

## What the review found

The corpus contained enough material to support a serious professional and creative archive, but it did not establish one canonical résumé, portfolio, visual identity, or asset-selection decision.

Several families were promising: a conventional résumé and visual variants, a client-facing technical plan, website fieldbooks and implementations, coordinated brand experiments, AlexOS visual explanations, governed trial records, and a large music-production archive.

Each carried a different evidence gap. Résumé claims needed corroboration. The technical plan needed permission and outcome evidence. Generated-image provenance was incomplete. Website lineage did not establish deployment authority. Music-session volume did not prove releases or finished works.

## Why restraint improved the result

Flattening the archive would have produced a cleaner folder and weaker evidence.

Preserving duplicates exposed shipment and preservation history. Keeping version families unresolved prevented filenames from becoming governance. Excluding sensitive material demonstrated that discovery could be useful without reading everything. Recording missing permission and provenance converted ambiguity into an actionable review queue.

The resulting principle is simple:

> Discovery may identify a candidate. It cannot grant the candidate authority.

## What this demonstrates

The inspected evidence supports these scoped claims:

- a read-only inventory covered an 8,474-file working corpus;
- bounded hashing generated a reproducible 983-file manifest and 79 exact-duplicate groups;
- copies, likely versions, provenance, sensitivity, and authority were treated as separate dimensions;
- source files were checked for size and modification-time stability after package creation;
- protected categories were deliberately excluded; and
- candidate selection produced no publication, canonicalization, or external action.

## Remaining limits

The review did not establish:

- semantic uniqueness across office documents;
- perceptual duplicate families;
- complete creative-source lineage;
- factual accuracy of résumé claims;
- client permission or project outcomes;
- release status of audio projects;
- ownership or usage rights for every asset; or
- a final portfolio or brand selection.

Those are separate investigations and human decisions. The discovery package made them visible without pretending to settle them.

