Commit Graph
17 Commits
Author SHA1 Message Date
etwenandClaude Opus 5 1a06367ade docs: Record V1.1.7 in Script A's history, and update CLAUDE / ARCHITECTURE
Script A's Version History gains the reboot-cycle macro, the Script B timeout
fix, the tmp/ rule and the stress knobs.

CLAUDE.md: current status to V1.1.7, a Script 5 paragraph, tmp/ in the folder
list, and two gotchas rewritten - the timeout one now covers both restore
positions, and --runtime is described as the job's overall limit rather than
"a 24 hour ceiling". Dropped the stale "TR518 not scripted yet" line.

ARCHITECTURE.md: Script 5 in the tree and in Key Features, tmp/ in the tree,
constraint 14 extended with "copy the :label along with the goto", and two
known issues added for the parts not yet run on hardware.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-09-04 22:44:21 +08:00
etwenandClaude Opus 5 78a8f641c9 fix(scriptc): Capture the BMC USB service log to a file, not the console
journalctl now runs with -o short-iso into log/bmc_usb_net.log and is
copied into the job directory with the other monitor logs. rm -f first so
the file covers this run only, and the ISO timestamps line up with
mgmt_ping.log and nfc_poll.log for cross-referencing.

The cost, recorded rather than glossed over: it is no longer echoed to
the console, so the USB copy is the only copy. That cp return value is
not checked - the same shape as the nfc_polling.log filename bug, which
failed silently and was only noticed back at the desk with a job archive
missing a file.

Folded into V1.1.6 rather than bumping: the tag has not moved on and this
is the same delivery.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-09-02 21:44:06 +08:00
etwenandClaude Opus 5 25e1ce0068 fix(pdb): Read the bricks on the pages the datasheet actually specifies
Every PDB reading came back NA, and the cause was a base error. The four
ADPM12200 bricks put each measurement on its own PMBus PAGE, written
before the read. The bring-up spreadsheet listed those page numbers in
decimal but with an 0x prefix:

    Vin  PAGE 9  -> it said 0x00   (0x09)
    Iin  PAGE 10 -> it said 0x10   (0x0A)
    Vout PAGE 2  -> it said 0x02    correct by luck
    Iout PAGE 14 -> it said 0x14   (0x0E)
    Temp PAGE 18 -> it said 0x18   (0x12)

Only Vout worked, because 2 reads the same in either base. Hence Iin and
Iout answering 0xFFFF, Temp answering 0x0000, and a Vin scale factor
having to be invented to make 50 V appear out of a page-0 register.

pwr_brick.sh reads all five bricks with the right pages and the
datasheet's DIRECT equation, X = (1/m)(Y x 10^-R - b): voltage Y x 8 mV,
current Y x 0.04 A, temperature Y x 0.01 C. Cross-checked rather than
assumed - READ_VOUT 0x05D4 decodes to 11936 mV, exactly what
'show platform voltage' reports for that rail through a separate sensor
path. The datasheet is committed alongside so the numbers are checkable.

Left flagged: Table 3 does not list READ_VIN or READ_IIN. They borrow the
voltage and current coefficients here, which is consistent with a ~50 V
input but is not something the datasheet states.

Each value is re-read until it passes three checks: not 0xFFFF, not
0x0000, and inside a plausibility window. The window is the one that
matters - this bus corrupts the HIGH BYTE only, which turns 11.98 V into
32.00 V. Positive, plausible in magnitude, and invisible to any all-ones
filter. 0x0000 is rejected for the mirror-image reason: on the
temperature register it decodes to a believable 0 C.

Adds comboA/comboB margin profiles per board. comboB is comboA with every
direction flipped, so the pair covers each rail high and low while its
neighbours sit the other way. NC channels stay nominal in both - there is
nothing to margin on an unused rail.

Script C re-enables the BMC USB journalctl dump, which had been commented
out, so that service log now reaches the master log.

Script A -> V1.1.6.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-09-02 16:35:45 +08:00
etwenandClaude Opus 5 215fed6e95 fix(stress): Keep the SSD and USB jobs alive for the whole soak
qfx5252-stress-usb was started with --runtime 60. It finished a minute
into an overnight run and the USB path sat idle for the rest of the
night - nothing failed, nothing said so, and the log looked complete.
SSD had the same shape at --runtime 3600. Both are now 86400, and the
BMC DDR memtester drops its loop count so it runs until stopped.

86400 is 24h, not forever. Past that SSD and USB die quietly and the
only clue is the shorter bgctl list Script B already prints each round.
Said so in the version history, the release notes and ARCHITECTURE.

Adds blanton_multiphase_margin.sh: the multiphase controllers on both
switch boards over the native i2c buses (SWB0 on i2c-13/15, SWB1 on
-14/-16), 19 rails each. high3/low3/comboA/comboB set voltages and print
one line per rail; save commits them and prints Multiphase Save All,
because that list has no per-rail mapping to report against.

Two hazards it carries: save is STORE_USER_ALL and permanent, and these
are NOT the LTC2980s margin.sh drives - the two tools reach the same
rails from opposite sides with no interlock between them.

The emitted commands were diffed against tools/Multiphase_*.txt: argv
identical for all five. The rail-to-netname mapping was cross-checked
against the bus numbers rather than trusted by line position - blocks
1..19 are all bus 13/15 and 20..38 all 14/16, matching the SWB0/SWB1
split of Multiphase_Netname.

blanton_traffic_linespeed v0.7.0 appends PatternRandom=yes to every tx
including ce0, and the defaults become tx 200 length=324. A fixed
pattern can sit on a benign bit sequence for a whole run; a changing one
is what exposes a weak lane. -r no forces fixed, -r "" omits the
argument for an SDK that will not take it. Bytes per burst went 51200 ->
64800, so absolute counter totals do not carry across this build - the
pair cross-check that decides PASS/FAIL does.

Script A -> V1.1.5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-09-01 11:04:45 +08:00
etwenandClaude Opus 5 db44f72352 feat(thermal): Add a standalone thermal/safety run, and move margin out of the soak loop
4_Blanton_Script_thermal_safety.ttl brings the DUT up the way Script A
does, then loads it - mlucas-avx2 -cpu 0:15 in the background plus
blanton_tr518.sh start - and samples platform data and both switch
boards' rail power every minute. Power per rail under sustained full
load, over hours, is the whole point of it.

It is standalone: nothing starts it, it starts nothing, and it must not
share a DUT with Script B because both take over the cd ports.

It also has no teardown. When you interrupt it the unit is still in
port cd lb=mac / l2 learn off / test mode nr=yes and mlucas is still
running, and show interfaces status does not reveal any of that. Called
out in the version history, the release notes, CLAUDE.md gotchas and the
ARCHITECTURE future-work list, because whoever uses the box next will not
find out on their own.

Script B's per-round loop drops the margin scan. The loop is meant to be
a quick minute-by-minute sample and margin_status_all.sh walks nine
LTC2980s over I2C, which dominated the round. Script C takes it instead,
once, before show uptime. The trade-off is real: margin drift DURING the
soak is no longer visible, only a start and an end reading. Documented
rather than glossed over, with a suggestion (scan every N rounds) if it
needs to come back.

Script C also sends exit and waits for "Script done" before
hw-test-session finish, so finish runs in the login shell rather than
inside the session it is closing.

Adds docs/Blanton_Test_Flow.drawio: all four scripts side by side, with
the config flags that change each step marked, Script B's one real goto
drawn, and the A->B->C handoffs shown as the manual steps they are.

The new script arrived LF-only while every other .ttl is CRLF and
.gitattributes marks *.ttl as -text, so what is committed is exactly what
Tera Term reads. Converted.

Script A -> V1.1.4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-31 22:19:38 +08:00
etwenandClaude Opus 5 9cc41f6fbd fix(margin): Reference VNOM to what the rails run at, and split profiles per CONN
86 of the 144 margin channels moved up +0.8% .. +2.5% to the measured
median for this platform (3.3 -> 3.38, 0.9 -> 0.915, 0.8 -> 0.82, ...).

That is not a display change. VNOM is the denominator of margin_status'
deviation %, and it is also the basis margin_set derives high/low/OV/UV
from - so both the number you read and the voltage you apply move with
it. A rail steady at 0.756 V used to read +0.80% against a nominal 0.75
and now reads 0.00%. Margin logs from before and after cannot be
compared: the reference moved, the hardware did not.

comboA/comboB are replaced by combo_SWB_CONN13..16_{high,low}. The
per-channel percentages are positional, so they follow the rails of one
specific board - applying a CONN13 profile to a CONN16 margins the wrong
rails and nothing detects it. SWB0 and SWB1 share a profile per CONN
because their rails are identical.

Fixes two things the deletion broke:
  - margin_apply_profile_all.sh still named comboA for all nine boards,
    which would now stop at "Profile not found" nine times over. It names
    the matching per-CONN profile instead, with CB on the generic
    combo_high3, and a note that the low pass is _high -> _low.
  - docs/LTC2980_channel_map.csv is generated from settings/*.conf, so
    all 86 changed rails were stale in it. Regenerated.

The eight new profiles also still carried comboA's header comment
("alternating high/low"), which describes neither what they contain nor
what they are for.

Script A -> V1.1.3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-30 22:13:43 +08:00
etwenandClaude Opus 5 bc2b529c3a feat(nfc): Poll the NFC tag through the soak instead of sampling it once
nfc_polling.sh wraps blantons_nfc_validate.py in a background loop and
keeps score: start/stop/status/tail/report/clear/fg. Script A takes a
30s baseline, Script B starts the poll for the soak, Script C stops it
with the other monitors and reports.

report only reads the log, so it can be run mid-soak without disturbing
anything - and it says RUNNING so a snapshot is not read as the final
answer. Each round appends a machine-readable RESULT line, which keeps
the tally independent of the tool's wording.

ERROR is counted apart from FAIL. A round where the tool printed neither
[PASS] nor [FAIL] - crashed, missing, hung past the timeout - measured
nothing, and that is a different fault from a tag that would not read.
Rolling them together hides whichever one you are not looking for.

Two habits from earlier bugs are baked in: tail does not follow (tail -f
never returns and a macro waiting on the prompt hangs there), and the .py
is called through python3 rather than ./ (it arrives mode 100644 and a
USB copy carries no execute bit).

Also fixes Script C copying "nfc_polling.log" to the job dir when the
file is "nfc_poll.log". The cp failed, nothing checked it, and the run
finished looking fine - you would find out at the desk, with a job
archive that had no NFC log in it.

Script A -> V1.1.2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 22:41:02 +08:00
etwenandClaude Opus 5 c3b0d49145 fix(session): Clear the last run's jobs and session before starting
After a run that ended badly - Tera Term closed, DUT rebooted mid-soak,
macro stopped by hand - the stress jobs kept running and the test session
stayed open. The next run then started on a box that was already loaded.
Nothing errored; the numbers were just quietly wrong.

Script A now opens with hw-test-session finish, bgctl reset --yes,
bgctl stop --all, then bgctl list. The list is the point of the sequence:
it is the one line in the log that proves the DUT was idle at the start.

Script A -> V1.1.1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 22:04:31 +08:00
etwenandClaude Opus 5 59861cf953 feat(ber,power): Add the BER sweep and TR518 load; keep ce0 actually sending
Three new tools and one fix to an old one.

blanton_ber.sh - PRBS BER across every cabled port, init/start/report/
stop/clear, -u 0|1|all. A lane passes only below 1e-6 (equal to the
threshold fails). report ends with a per-unit summary carrying the worst
lane, so an all-PASS run still shows how much margin there was. init
checks every lane locked, because an unlocked lane still reports a BER
and it is a meaningless one. Script C runs it after the traffic report,
following SWB_UNIT0/SWB_UNIT1 like the traffic block does.

blanton_tr518.sh - the built-in packet test, used to load the rails
before measuring power. start settles 120s so the reading afterwards is
loaded rather than mid-ramp. stop sends port cd lb=none and nothing else,
exactly what the bench procedure does - l2 learning and test mode are
left as start set them, and stop says so instead of quietly leaving the
box in a state nobody asked about.

TH6_SWB{0,1}_power_readback.sh - 19 rails, V x I and total. These read
the two sensor tables ONCE. The previous version read them per rail: 38
invocations of a slow CLI at 38 different instants, so the "total" was a
sum of readings seconds apart - and under load the currents move (TH6_CORE
186 A -> 636 A). A rail missing from the table now warns on stderr instead
of silently contributing 0 W.

blanton_traffic_linespeed.sh - ce0 (VLAN 138) joins the test. It has no
loopback partner and a switch never sends a frame back out its ingress
port, so a tx burst was one round trip and then silence. init now sets an
ingress mirror of ce0 back to itself; the mirrored copy is not
egress-filtered, so packets keep lapping until stop turns the mirror off.
stop is what ends it - removing the VLAN membership does not, because the
copy never consulted the VLAN. TL_CE_MODE also drops to 'self': ps ce
shows one ce port per unit and a unit-0-only burst returned on unit 0.
Read as cross-unit, both links dying would still report PASS (0 == 0).

Fixes a parser bug worth remembering: bcmcmd returns CRLF, and the BER is
the last token on its line, so the CR stayed glued to the number. Every
lane read as NA, and printing the CR sent the cursor to column 1 so the
RESULT column overwrote the start of the row. One stray \r, two symptoms.

Not yet verified on hardware, all called out in the release notes: the
ce0 mirror (destport == srcport may be refused), the 120s settle (tr 518
may be time-limited), and 'test mode nr=no' on stop.

Script A -> V1.1.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 16:33:40 +08:00
etwenandClaude Opus 5 43314444e9 feat(config): Run the fans at full speed by default
FAN_SPEED 30 -> 100. A soak is a thermal test of everything except the
cooling, so the cooling should not be one of the variables in it. At 30%
a long run can throttle partway through, and throttled results read like
a different fault entirely.

Folded into V1.0.11 rather than bumping the version: the tag has not
moved on and this is the same delivery.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 10:32:07 +08:00
etwenandClaude Opus 5 4854da8518 feat(margin): Save every LTC2977 in one command, and stop lying about failures
margin_save discarded the result of every write, so a NACK printed
"Change Saved" exactly like a successful store. For the one permanent
operation in this tool that is the worst place to stay quiet. Each store
is now checked, named on failure, and reflected in the exit code.

STORE_USER_ALL also went out as 0x15 followed by a dummy 0x00 - an SMBus
write-byte where the command is a send-byte. The LTC2977 tolerated it,
but tolerance is not correctness. Now i2cset ... 0x15 c on the native
path and a no-data cb_pmbus_write on the FPGA path.

Two new subcommands:

  margin_save all      - all nine boards in settings/, then re-init
                         whichever board was loaded beforehand so a
                         following margin_status still reports the board
                         you were looking at. Does not stop on the first
                         failure; a half-saved set is harder to reason
                         about than a fully-attempted one.

  margin_save blanton  - this platform exposes the CB FPGA F3 I2C
                         channels as native Linux i2c buses (Ch6->13,
                         Ch7->14, Ch8->15, Ch9->16; CB stays on bus 4),
                         so all 18 LTC2977 can be stored with plain
                         i2cset: no margin_init, no FPGA channel setup,
                         no pcimem. 18 commands instead of 42.

The 18 raw commands were verified on the DUT 2026-08-27; the wrapper was
tested against stubbed hardware only. Release notes say so.

Board list is MARGIN_BLANTON_STORE at the top of margin.sh - comment out
what this DUT does not have rather than letting it talk to absent boards.
Missing bus and unresponsive chip are reported as distinct failures.

MARGIN_STORE_SETTLE (0.5s) sits between stores because nothing here polls
the part's busy bit, and blanton fires 18 in a row.

Script A -> V1.0.11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 08:39:31 +08:00
etwenandClaude Opus 5 15ea776b82 feat(soak): Show the background jobs each round, and stop the loop spinning
Script B's monitoring loop had no pause at all - each round started the
instant the last one ended, so an overnight soak filled the log with
near-identical samples taken seconds apart. The section was even labelled
"Get data every 10mins" while running with no delay whatsoever.

The loop now ends each round with pause 60 and prints both job lists:
bgctl list for the platform job runner, jobs for anything the macro
backgrounded in that shell. A stress job that dies mid-soak shows up as a
shorter list on the next round instead of going unnoticed until the end.

60s is still hard-coded and the Status sheet asks for 10 minutes, so the
Phase 3 item stays open - this makes the loop sane, it does not close it.

Also adds docs/LTC2980_channel_map.csv: the nine settings/*.conf flattened
into one Board,CONN,Ch,NetName,Vnom sheet, generated by
tools/gen_channel_map.sh. The CSV is output, not source. NC channels are
kept as NC/- because the channel number is also the PMBus page - dropping
the gaps would shift every channel after them.

Two things the merged table makes visible: SWB0 and SWB1 carry identical
net names and voltages (only the bus and CB_I2C_CH differ), so those are
two copies that can drift apart silently; and there are exactly 8 NC
channels, all on CONN14/CONN15.

tools/100G_PRBS.txt, TR518.txt and fan_ctrl.txt join the existing
traffic_loopback_*.txt references. TR518 is not scripted yet and cannot
share the hardware with blanton_traffic_linespeed - both drive port
loopback and L2 learning.

Script A -> V1.0.10.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-27 14:29:47 +08:00
etwenandClaude Opus 5 03d609cc1f feat(prbs): Add 100G PRBS monitor; fix the job log path and USB mount
port_prbs_monitor.sh drives PRBS on both 100G uplinks: lpmode off, arm
the pattern, then poll `phy diag <phy> prbs get` in the background until
stopped, counting a poll as PASS only when the output says "PRBS OK!".
prbsstat Ber is captured beside every poll for the record but does not
decide the verdict. `stop` runs prbsstat STOp and prbs clear before
rendering the report, so the test does not leave PRBS armed.

`start a` / `start b` runs one uplink alone, which is how a genuine port
fault is separated from the DUT not coping with two PRBS streams at
once. A port that was not run reports SKIP rather than FAIL -- deciding
on poll count alone would have made a deliberate single-port run look
like half the hardware was broken.

The port status is deliberately not printed at `start`: with PRBS armed
the link reports DOWN, which reads as a failure to anyone glancing at
the console. It appears in the report instead, after prbs clear and
labelled as the recovered state.

The calls in Script A, B and C are committed but commented out -- PRBS
is still under bring-up.

Fix: the job directory was written as /host/hw-eval/current/jobs in some
places and /host/hw-eval/jobs in others, so what Script B wrote was not
what Script C collected. Unified on /host/hw-eval/jobs.

Fix: Script C copied to /mnt/usb assuming it was mounted. It now finds
the device with usb_target.sh and mounts it first, so the logs actually
leave the DUT.

Script C also reorders its results -- stress logs, then the MGMT ping
report, then 100G status -- and cats the bgctl ssd/usb logs.

Script A -> V1.0.9. Docs and docs/release-notes/v1.0.9.md updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-26 15:35:48 +08:00
etwenandClaude Opus 5 5e1e4a6a8f docs(release): Add v1.0.8 notes; wire mgmt ping status into B and C
Script A: the baseline ping window is 30 s rather than 45, and the log
is cleared afterwards so Script B starts from nothing.

Script B: `status` after `start`, and Script C: `status` before `stop`
plus a `cat` of the report. The status line shows both pids and the
replies received per NIC, so a dead link is visible while the soak is
still running instead of only at the end.

docs/release-notes/v1.0.8.md written in the house format.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-26 10:13:34 +08:00
etwenandClaude Opus 5 ef7cbdffd4 docs(release): Add v1.0.7 notes; bump Script A header to V1.0.7
Script A carried a V1.0.7 history block while its "; Version :" header
still read V1.0.6, so publish.sh built Script_ABC_Blanton_V1.0.6 and the
folder would have disagreed with the tag. Header and date corrected --
reading the version from the header rather than a flag is what surfaced
this.

utils/wait_init.ttl V3.1.1: poll interval 10 s -> 60 s, and the WT_MIN
comments now say "at least" to match the test, which has been
`< WT_MIN` since the threshold was fixed. The stale "more than" wording
described exactly the off-by-one that once hung Script A.

Script C: the BMC USB journalctl dump is commented out.

docs/release-notes/v1.0.7.md written in the house format.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 14:22:22 +08:00
etwen 9d61b1aa43 docs(status): Add TTL script status sheet 20260817 2026-08-17 16:32:44 +08:00
etwenandClaude Opus 5 24c704c667 docs(arch): Add ARCHITECTURE.md and CLAUDE.md, init repo
Bring the Blanton Tera Term TTL test suite under version control and
document how the pieces fit together.

- ARCHITECTURE.md: A/B/C script roles, host-to-DUT-to-FPGA data path,
  config/settings/profile data models, PCIe AER + EDAC topology table,
  10 key constraints, and 6 development phases derived from the
  Status_20260814 spreadsheet's On-Going items
- CLAUDE.md: stack, smoke-test commands, conventions, and the hardware
  footguns (per-unit BDF, VSPI vs VI2C timing, repeated START,
  STORE_USER_ALL, ARB_LOST false positives)
- secret/: gitignored credential store with README + .example template,
  so the DUT login stops living in 1_Blanton_Script_A.ttl
- .gitignore: secret/*, For_AI/, *.log, publish/, *.DSN

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-17 08:52:11 +08:00