21 Commits
Author SHA1 Message Date
etwenandClaude Opus 5 43314444e9 feat(config): Run the fans at full speed by default
FAN_SPEED 30 -> 100. A soak is a thermal test of everything except the
cooling, so the cooling should not be one of the variables in it. At 30%
a long run can throttle partway through, and throttled results read like
a different fault entirely.

Folded into V1.0.11 rather than bumping the version: the tag has not
moved on and this is the same delivery.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 10:32:07 +08:00
etwenandClaude Opus 5 4854da8518 feat(margin): Save every LTC2977 in one command, and stop lying about failures
margin_save discarded the result of every write, so a NACK printed
"Change Saved" exactly like a successful store. For the one permanent
operation in this tool that is the worst place to stay quiet. Each store
is now checked, named on failure, and reflected in the exit code.

STORE_USER_ALL also went out as 0x15 followed by a dummy 0x00 - an SMBus
write-byte where the command is a send-byte. The LTC2977 tolerated it,
but tolerance is not correctness. Now i2cset ... 0x15 c on the native
path and a no-data cb_pmbus_write on the FPGA path.

Two new subcommands:

  margin_save all      - all nine boards in settings/, then re-init
                         whichever board was loaded beforehand so a
                         following margin_status still reports the board
                         you were looking at. Does not stop on the first
                         failure; a half-saved set is harder to reason
                         about than a fully-attempted one.

  margin_save blanton  - this platform exposes the CB FPGA F3 I2C
                         channels as native Linux i2c buses (Ch6->13,
                         Ch7->14, Ch8->15, Ch9->16; CB stays on bus 4),
                         so all 18 LTC2977 can be stored with plain
                         i2cset: no margin_init, no FPGA channel setup,
                         no pcimem. 18 commands instead of 42.

The 18 raw commands were verified on the DUT 2026-08-27; the wrapper was
tested against stubbed hardware only. Release notes say so.

Board list is MARGIN_BLANTON_STORE at the top of margin.sh - comment out
what this DUT does not have rather than letting it talk to absent boards.
Missing bus and unresponsive chip are reported as distinct failures.

MARGIN_STORE_SETTLE (0.5s) sits between stores because nothing here polls
the part's busy bit, and blanton fires 18 in a row.

Script A -> V1.0.11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-28 08:39:31 +08:00
etwenandClaude Opus 5 15ea776b82 feat(soak): Show the background jobs each round, and stop the loop spinning
Script B's monitoring loop had no pause at all - each round started the
instant the last one ended, so an overnight soak filled the log with
near-identical samples taken seconds apart. The section was even labelled
"Get data every 10mins" while running with no delay whatsoever.

The loop now ends each round with pause 60 and prints both job lists:
bgctl list for the platform job runner, jobs for anything the macro
backgrounded in that shell. A stress job that dies mid-soak shows up as a
shorter list on the next round instead of going unnoticed until the end.

60s is still hard-coded and the Status sheet asks for 10 minutes, so the
Phase 3 item stays open - this makes the loop sane, it does not close it.

Also adds docs/LTC2980_channel_map.csv: the nine settings/*.conf flattened
into one Board,CONN,Ch,NetName,Vnom sheet, generated by
tools/gen_channel_map.sh. The CSV is output, not source. NC channels are
kept as NC/- because the channel number is also the PMBus page - dropping
the gaps would shift every channel after them.

Two things the merged table makes visible: SWB0 and SWB1 carry identical
net names and voltages (only the bus and CB_I2C_CH differ), so those are
two copies that can drift apart silently; and there are exactly 8 NC
channels, all on CONN14/CONN15.

tools/100G_PRBS.txt, TR518.txt and fan_ctrl.txt join the existing
traffic_loopback_*.txt references. TR518 is not scripted yet and cannot
share the hardware with blanton_traffic_linespeed - both drive port
loopback and L2 learning.

Script A -> V1.0.10.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-27 14:29:47 +08:00
etwenandClaude Opus 5 03d609cc1f feat(prbs): Add 100G PRBS monitor; fix the job log path and USB mount
port_prbs_monitor.sh drives PRBS on both 100G uplinks: lpmode off, arm
the pattern, then poll `phy diag <phy> prbs get` in the background until
stopped, counting a poll as PASS only when the output says "PRBS OK!".
prbsstat Ber is captured beside every poll for the record but does not
decide the verdict. `stop` runs prbsstat STOp and prbs clear before
rendering the report, so the test does not leave PRBS armed.

`start a` / `start b` runs one uplink alone, which is how a genuine port
fault is separated from the DUT not coping with two PRBS streams at
once. A port that was not run reports SKIP rather than FAIL -- deciding
on poll count alone would have made a deliberate single-port run look
like half the hardware was broken.

The port status is deliberately not printed at `start`: with PRBS armed
the link reports DOWN, which reads as a failure to anyone glancing at
the console. It appears in the report instead, after prbs clear and
labelled as the recovered state.

The calls in Script A, B and C are committed but commented out -- PRBS
is still under bring-up.

Fix: the job directory was written as /host/hw-eval/current/jobs in some
places and /host/hw-eval/jobs in others, so what Script B wrote was not
what Script C collected. Unified on /host/hw-eval/jobs.

Fix: Script C copied to /mnt/usb assuming it was mounted. It now finds
the device with usb_target.sh and mounts it first, so the logs actually
leave the DUT.

Script C also reorders its results -- stress logs, then the MGMT ping
report, then 100G status -- and cats the bgctl ssd/usb logs.

Script A -> V1.0.9. Docs and docs/release-notes/v1.0.9.md updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-26 15:35:48 +08:00
etwenandClaude Opus 5 5e1e4a6a8f docs(release): Add v1.0.8 notes; wire mgmt ping status into B and C
Script A: the baseline ping window is 30 s rather than 45, and the log
is cleared afterwards so Script B starts from nothing.

Script B: `status` after `start`, and Script C: `status` before `stop`
plus a `cat` of the report. The status line shows both pids and the
replies received per NIC, so a dead link is visible while the soak is
still running instead of only at the end.

docs/release-notes/v1.0.8.md written in the house format.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-26 10:13:34 +08:00
etwenandClaude Opus 5 a31c186142 feat(mgmt): Ping both NICs at once, and stop counting ARP as loss
The bench run reported FAIL=3 with no NIC errors and no drops. All three
failures were on eth1 and all three were missing exactly icmp_seq=1,
with the other nine replies present and sub-millisecond. That is
neighbour resolution: once the ARP entry for the target expires, the
first echo request is spent resolving it and ping counts it as loss.
Rounds sat about 24 s apart, right on the edge of the default reachable
time, so it happened on some rounds and not others.

mgmt_ping_monitor.sh V3.0.0:

- One discarded ping per NIC before the measured run resolves ARP, so
  MAX_LOSS_PCT can stay at 0 and a FAIL means a real drop. Raising the
  threshold instead would have hidden genuine single-packet loss, and a
  longer settle would not help -- nothing triggers ARP until the ping.
- Both NICs now ping at the same time and keep accumulating until
  stopped, rather than alternating fixed bursts. Each writes its own raw
  capture.
- `stop` sends SIGINT rather than SIGTERM, because ping prints its
  statistics block on interrupt and that block is what the report
  parses, then renders into the log: the last TAIL_LINES entries per
  NIC, each one's statistics and verdict, and `ip -s link show` for both
  interfaces.
- `-D -O` are on by default so every line carries a timestamp and a
  request that got no reply is visible in the tail instead of merely
  absent. Clear PING_EXTRA_OPTS if the platform's ping rejects them.

Since the script no longer re-execs itself to daemonise, `bash
mgmt_ping_monitor.sh start` works without the execute bit.

.gitignore: the new *.raw / *.pid / *.state runtime files were outside
the *.log rule and would have been committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-26 09:53:20 +08:00
etwenandClaude Opus 5 1b8d5a8e91 feat(usb): Detect the USB target instead of assuming /dev/sda1
New Blanton_Script/usb_target.sh reports a USB mass-storage device's
node, mount point or by-id path. It excludes whatever disk backs / and
/host -- this platform can boot from a USB DOM, which also reports
TRAN=usb -- refuses to guess when several USB disks are present, falls
back to the whole device for superfloppy media, and when mounting,
checks the result is actually writable rather than trusting that mount
succeeded.

Script B asks it for the device node and builds the stress command from
the answer, so a stick that enumerates as sdb no longer sends the write
test at whatever /dev/sda1 happens to be.

The device is carried back through an echo marker, split in the shell
literal as "<<D""EV=..." so the command echo cannot satisfy the pattern
the macro waits for -- the same false-match guard used for PORTS0= and
NCMIF=, arranged differently.

Script B: DDR memtester drops its 100-pass limit and runs until stopped,
matching the other soak loads.

Fix: Script C ran `bgctl reset --yes` while killing processes, which is
before the job logs are copied to USB. Moved to after the archive and
after hw-test-session finish, so the run's own logs are collected before
anything clears them.

Script A -> V1.0.8 history updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 17:03:36 +08:00
etwenandClaude Opus 5 cac37583e7 feat(ttl): Include mgmt_ping.log in the USB archive
Script C now copies log/mgmt_ping.log into
/host/hw-eval/current/jobs/ alongside bmc_poll.log, so the management
link results leave the DUT with the run instead of staying behind. Both
monitors' logs are now covered.

Script A history: the two copies are one action in one block, so they
share a single V1.0.8 line rather than reading as separate changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 15:57:23 +08:00
etwenandClaude Opus 5 802ee4d467 feat(ttl): Fold bmc_poll.log into the USB archive
Script C copies log/bmc_poll.log into /host/hw-eval/current/jobs/ before
that directory is archived to USB, so the BMC sampling and I2C
read-back results leave the DUT with the run they belong to rather than
only existing in the Tera Term capture.

The wait moves ahead of the getdate/sprintf2 block; those build strings
without touching the terminal, so the prompt is still consumed exactly
once before the copy is sent.

Script A history updated under the existing V1.0.8 block -- that version
is not tagged yet, so this is the same batch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 15:55:05 +08:00
etwenandClaude Opus 5 166b53983f feat(bmc): Re-enable bmc_monitor with an I2C write/read-back test
bmc_monitor.sh comes back into the run with a second job beyond the
`free -m` sampling: an I2C pattern test against bus 1, address 0x41,
offset 0x00C0 -- write 55AA55AA and read it back, then write AA55AA55
and read that back. Two complementary patterns catch a stuck bit either
way round, and repeating it through the soak turns an intermittent I2C
fault into something the log shows rather than something the tester has
to catch live.

Script B starts it alongside the other stress; Script C stops it and
cats log/bmc_poll.log.

bmc_monitor_ddr.sh stays disabled -- BMC DDR stress runs through
bgctl + bmc-manager.

Docs: the "disabled" notes I added last time now applied to only half
the pair and read as wrong for bmc_monitor.sh. Split so the monitor is
documented as live with its I2C test, and only the _ddr variant is
marked disabled.

Script A -> V1.0.8.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 15:50:56 +08:00
etwenandClaude Opus 5 ef7cbdffd4 docs(release): Add v1.0.7 notes; bump Script A header to V1.0.7
Script A carried a V1.0.7 history block while its "; Version :" header
still read V1.0.6, so publish.sh built Script_ABC_Blanton_V1.0.6 and the
folder would have disagreed with the tag. Header and date corrected --
reading the version from the header rather than a flag is what surfaced
this.

utils/wait_init.ttl V3.1.1: poll interval 10 s -> 60 s, and the WT_MIN
comments now say "at least" to match the test, which has been
`< WT_MIN` since the threshold was fixed. The stale "more than" wording
described exactly the off-by-one that once hung Script A.

Script C: the BMC USB journalctl dump is commented out.

docs/release-notes/v1.0.7.md written in the house format.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 14:22:22 +08:00
etwenandClaude Opus 5 ae8941a844 feat(ttl): Set fan speed in Script A, archive job logs to USB in C
config.ttl gains FAN_SPEED (30). Script A rebinds the two max31790
controllers (18-0020, 25-0020) and applies the speed with
fan-speed-control.sh, answering its confirmation prompt, so a run starts
from a known thermal state rather than whatever the last test left.

Script A also clears /host/hw-eval/current/jobs before the run, and
Script C copies that directory to /mnt/usb/jobs-<date>_<time> so the
bgctl job logs leave the DUT with the run they belong to.

Fix: 25-0020 was bound twice in a row -- the second bind can only fail,
since the driver is already attached. Removed; the two controllers now
have a matching unbind/sleep/bind each.

Script A -> V1.0.7, history covering the config.ttl and Script C changes
as well.

Docs: fan control and the job-log archive added to Tech Stack, Key
Features and Data Flow; the config.ttl example refreshed with FAN_SPEED
and the SWB_UNIT flags. Three risks recorded in Future Extensions --
`rm` without -r cannot clear the job subdirectories that Script C copies
with -r, the fan confirmation wait has no timeout and would hang
silently if the prompt ever changes, and /mnt/usb is never checked for
being mounted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-25 14:12:24 +08:00
etwenandClaude Opus 5 a6a12cdf2c docs: Sync ARCHITECTURE and CLAUDE with the bgctl and mgmt-ping work
Script A: wait 45 s between starting and stopping mgmt_ping_monitor.
`start` returns after one second, but a round is roughly 30 s -- without
the pause the log held only the START banner and no RESULT line.

Docs brought up to date with the eight commits since 419b680:

- Tech Stack gains bgctl, systemd-run, iproute2 and hw-test-session;
  the stress row moves off the hammer scripts
- Structure and Key Features cover mgmt_ping_monitor.sh, and mark
  bmc_monitor{,_ddr}.sh as disabled -- their call sites in B and C are
  commented out, though the files remain
- Data Flow rewritten for all three scripts
- Two constraints added: TTL's `timeout` is global and must be restored,
  and two management NICs on one subnet cause ARP flux (hence
  ARP_STRICT, and nic_tx/nic_rx as the evidence of which NIC sent)
- CLAUDE gotchas add the `ip -s link` column order, that a monitor's
  `start` returns in one second rather than after a round, and that
  show_dmesg now ends with `dmesg -C`
- Phase 3 ticks the move to bgctl; Phase 5 gains the mgmt-ping, NVMe and
  EDAC parsing targets. Future Extensions notes that Script C never cats
  mgmt_ping.log, so the soak's ping results stay on the DUT

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-24 23:34:33 +08:00
etwenandClaude Opus 5 e38d24c775 feat: Add mgmt_ping_monitor.sh and move stress onto bgctl
New Blanton_Script/mgmt_ping_monitor.sh: brings both management NICs up
with iproute2 (`ip link set` / `address replace` / `route replace` with
per-NIC metric) and pings each one's own target -- eth0 -> .30, eth1 ->
.31 -- appending to a rotating log. Same sub-commands as the bmc
monitors, plus `summary` for just the per-leg RESULT lines.

Both NICs stay up, which on a shared subnet lets the target's ARP be
answered by either one, so a reply can land on the NIC that did not
send. ARP_STRICT applies arp_ignore/arp_announce to prevent that, and
each leg logs `ip -s link show` with the interface's own TX/RX packet
delta across the burst as direct evidence of which NIC carried the
traffic. Note `ip -s link` orders columns "bytes packets ...", so the
packet count is the second field.

Script A: hw-test-session start/log/status, a fuller DUT inventory
(version, fwutil, syseeprom, ssdhealth, TPM, nvme smart-log, smartctl),
bmc-first-enroll and bmc-manager version/status, ras-mc-ctl summary, and
both 100G uplinks now brought up rather than 513 being left down.

Script B: stress moves to bgctl (memtester, qfx5252-stress-ssd/-usb) and
the BMC DDR load runs through bmc-manager, replacing the hammer scripts
and the bmc_monitor pair. Adds a BMC USB net test that discovers the
cdc_ncm interface and runs a 4-hour ping under systemd-run.

Script C: stops the bgctl jobs, the BMC USB unit and the mgmt ping, then
collects journalctl, per-NIC counters, ras-mc-ctl and NVMe health, and
closes the session with hw-test-session finish.

utils/show_dmesg.ttl: one combined error regex with `dmesg -T`, an
i2c-filtered view, then `dmesg -C` so Script C's capture shows only what
the soak produced. utils/setup_pmon.ttl: add the TPM FRU/read checks.

Fix: Script B set `timeout = 15` for the cdc_ncm probe and never
restored it, leaving the cap in force for everything after -- including
the traffic init, which walks 108 VLANs per unit and takes far longer
than 15 s. A timed-out wait returns without the prompt, so the macro
would have run ahead of the DUT for the rest of the soak. Reset to 0 at
:skip_ping, where both branches meet.

Script A -> V1.0.6.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-24 23:27:37 +08:00
etwenandClaude Opus 5 9a11c1d5bb fix(ttl): Wait for the prompt before config save in the uplink block
`config save -y` was sent straight after the Ethernet514 startup with no
wait in between. The shell buffers the second line and still runs it, so
nothing visibly breaks, but every wait from there on matches the prompt
of the previous command -- leaving the macro permanently one step ahead
and issuing `show interfaces status` while `config save` is still
running. Every other sendln in the file waits first; this one now does
too.

Version left at V1.0.5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 12:58:06 +08:00
etwenandClaude Opus 5 cbef5865ec fix(ttl): Treat 216 ports up as ready, and configure the 100G uplinks
wait_init.ttl: the comparison was strictly greater than WT_MIN, so a unit
reporting exactly 216 never passed and Script A sat in the poll loop
forever. 216 is not an arbitrary number -- TL_PAIRS holds 108 loopback
pairs, so 108 x 2 = 216 is every cabled port being up, i.e. precisely the
state being waited for. Comparison is now `< WT_MIN`, making the
threshold "at least 216".

Script A: bring the 100G uplinks into a known state before reading their
status -- Ethernet513 on asic0, Ethernet514 on asic1 -- and persist it
with `config save -y`. The stray `wait` after show_dmesg is dropped; the
one opening the new block consumes that prompt instead.

Version left at V1.0.5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 12:56:13 +08:00
etwenandClaude Opus 5 8714bbb773 fix(ttl): Compare SWB_UNIT1 in the second operand of the traffic branches
Rework the traffic blocks in A, B and C as explicit if/elseif branches
holding literal commands, and drop the derived swb_any/swb_opt from
config.ttl -- the branches read straight off the page and there is no
indirection to follow.

All nine conditions compared SWB_UNIT0 against itself, though:

    if     SWB_UNIT0 = 1 && SWB_UNIT0 = 1     ->  SWB_UNIT0 = 1
    elseif SWB_UNIT0 = 1 && SWB_UNIT0 = 0     ->  never
    elseif SWB_UNIT0 = 0 && SWB_UNIT0 = 1     ->  never

so the single-unit paths were unreachable. A DUT with only unit 0 would
have run the both-unit commands and hit bcmcmd on an absent unit 1,
while a DUT with only unit 1 fell through to the empty else and skipped
traffic entirely. Second operand is now SWB_UNIT1.

Verified per branch that the -u argument matches the condition guarding
it: both units carry no -u (tool default TL_UNITS="0 1"), unit-0-only
carries -u 0, unit-1-only carries -u 1, and the else stays empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 12:19:18 +08:00
etwenandClaude Opus 5 d8bf1f7a41 feat(ttl): Drive traffic and port wait from SWB_UNIT0/SWB_UNIT1
The bench is not always fully populated. config.ttl now declares which
switch units exist:

    SWB_UNIT0 = 1   ; this DUT has switch unit 0
    SWB_UNIT1 = 1   ; this DUT has switch unit 1

Two values are derived there rather than repeating the same test in
three scripts: swb_any (0 = no unit at all) and swb_opt, the suffix
appended to each blanton_traffic_linespeed call -- "" for both units so
the tool's own TL_UNITS="0 1" applies, " -u 0" or " -u 1" for a single
one.

Script A, B and C build their traffic commands with sprintf2 and the
suffix, wrapped in `if swb_any = 1`. A half-populated DUT no longer
issues bcmcmd against an absent unit, and a DUT with no switch board
skips the traffic stage outright instead of filling the log with
failures.

wait_init.ttl reads the same two flags instead of its own copies, so the
wait and the traffic blocks cannot disagree about what is installed.

Script A -> V1.0.5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 12:06:28 +08:00
etwenandClaude Opus 5 da39122c43 feat(ttl): Let wait_init skip switch units the DUT does not have
wait_init.ttl V3.0.0 adds WT_UNIT0 and WT_UNIT1 at the top of the file.
Benches are not always fully populated, and V2.0.0 waited on both units
unconditionally, so a DUT with one switch board sat in the poll loop
forever -- silently, since the loop neither advances nor reports.

    both units  -> 1 , 1
    unit 0 only -> 1 , 0
    unit 1 only -> 0 , 1
    no unit     -> 0 , 0   (bypass, returns immediately)

A disabled unit is skipped rather than polled and ignored: its whole
block sits inside the if, so no bcmcmd is issued for it and no
misleading error reaches the log.

Two independent integer flags rather than a "0 1" string, because
parsing a string in TTL needs strscan and this is meant to be edited by
hand at the bench.

All three exit paths still leave exactly one prompt unconsumed, so
Script A's surrounding waits are unaffected.

Script A history: recorded under the existing V1.0.4 block rather than a
new version, matching the consolidation done there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 11:36:38 +08:00
etwenandClaude Opus 5 b66ae47a58 feat(ttl): Read EEPROM, quiet LLDP and clear lpmode before traffic
Script A: `hpe-eeprom-tlv show --bus 4 --addr 0x50` after `show boot`, so
the board identity is on record next to the image it booted. The empty
Check History placeholder is dropped.

Script A and B: disable the lldp feature and `config save -y` before the
traffic stage, so the switch stops sourcing its own frames and the
loopback pair counters reflect only the injected burst.

Script B: `sfputil lpmode off` on Ethernet513/514 before reading their
status -- a transceiver left in low-power mode will not link.

Script A -> V1.0.5, with the history covering the Script B changes too,
per this project's convention of keeping one consolidated log in A.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 11:28:06 +08:00
etwenandClaude Opus 5 d065db7d71 feat(ttl): Gate Script A on bcmcmd port-up count, not thermal sensors
wait_init.ttl V2.0.0 now polls

    bcmcmd -n 0 -c ps | grep -w up | wc -l
    bcmcmd -n 1 -c ps | grep -w up | wc -l

and proceeds only when BOTH exceed 216, so the baseline is taken with
the data plane actually up rather than merely with pmon answering.

V1.0.0 could use `wait "Thermal Not detected" prompt` because that was a
string-presence test. Comparing a count needs the value captured, so the
count is wrapped in an echo marker and read with waitregex +
groupmatchstr1 + str2int. The command echo cannot false-match: it reads
"PORTS0=$(bcmcmd ..." and the pattern requires a digit immediately after
the "=". The optional-space allowance covers a wc that pads its output.

The two thresholds are compared in nested ifs rather than with `and`,
which is bitwise in TTL.

WT_MIN and WT_INTERVAL are at the top of the file. The enter/exit prompt
contract is unchanged, so Script A's surrounding waits still line up.
Script A -> V1.0.4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
2026-08-21 10:41:50 +08:00
28 changed files with 2836 additions and 160 deletions
+5
View File
@@ -12,6 +12,11 @@ For_AI/
!**/logs/.gitkeep
!**/Logs/.gitkeep
# monitor 腳本的執行期產物(原始擷取、pid、狀態)
*.raw
*.pid
*.state
# 打包輸出:產出不進 git,打包工具本身要進
publish/*
!publish/.gitkeep
+146 -39
View File
@@ -37,8 +37,13 @@ LTC2980 margin 控制器。TTL 只負責「按順序 `sendln` 並把畫面收進
| 電源 margin | LTC2980= 2 × LTC2977over PMBusLINEAR16 編碼 |
| 錯誤計數來源 | Linux PCIe AER sysfs`aer_dev_*`)、EDAC sysfs`dimm_ce/ue_count` |
| 平台監控 | SONiC `show platform *`pmon container,含 `leak status` / `leak channels` |
| BMC 存取 | host 端 `bmc-manager run "<cmd>"`(不需對 BMC 開 SSH,繞開其 busybox 工具鏈) |
| 壓力測試 | `mlucas-avx2`AMD CPU)、`~/hammer/tools/stress_{mem,ssd,usb}.py` |
| BMC 存取 | host 端 `bmc-manager run "<cmd>"`(不需對 BMC 開 SSH,繞開其 busybox 工具鏈)`bmc-first-enroll` / `bmc-manager version\|status` |
| 背景工作管理 | `bgctl run/list/stop --all/reset`(平台工具);長時間 ping 用 `systemd-run --unit=` 掛成 transient unit |
| 管理網路 | iproute2`ip link set` / `ip address replace` / `ip route replace ... metric`+ `ping -I` |
| 測試 session | `hw-test-session start\|log\|status\|finish`job log 收在 `/host/hw-eval/jobs/` |
| 風扇控制 | max31790 driver rebind`/sys/bus/i2c/drivers/max31790/{bind,unbind}`+ `fan-speed-control.sh <%>` |
| 100G PRBS | `bcmcmd -n <unit> -c 'dsh -c "phy diag <phy> prbs …"'`set / prbsstat STArt / get / Ber / STOp / clear |
| 壓力測試 | `mlucas-avx2`AMD CPU)、`bgctl run` 排程的 `memtester` / `qfx5252-stress-ssd` / `qfx5252-stress-usb` |
| 資料面流量 | `bcmcmd`Broadcom drivshellVLAN loopback + `tx` burstswitch unit 0 / 1 |
| 流量報表 | DUT 上 awk`blanton_traffic_linespeed report`);離線 `python3 tools/bcm_mibpair_report_V1.1.0.py` |
| 版本控制 | GitNAS + Gitea 私有 remote**不推 GitHub,含客戶 NDA 資料** |
@@ -115,6 +120,7 @@ Blanton_TTL_Script/
│ ├── Blantons_FPGA_Registers_draft.docx # 同上,客戶原始 docx
│ ├── Blantons_FPGA_Registers_Map_draft.xlsx
│ ├── Blantons_PCIe_AER_and_DDR_EDAC_checks.pdf # AER / EDAC 檢查點依據
│ ├── LTC2980_channel_map.csv # 9 個 settings/*.conf 合併:Board,CONN,Ch,NetName,Vnom144 列)
│ ├── TTL_Script_Blanton_Status_20260814.xlsx
│ └── TTL_Script_Blanton_Status_20260817.xlsx # ⭐ 最新:三支腳本逐項 StatusOK / On-Going + Owner
@@ -122,9 +128,13 @@ Blanton_TTL_Script/
│ ├── bcm_mibpair_report_V1.1.0.py # 離線解析 drivshell log → per-pair TX/RX + PASS/FAIL
│ ├── traffic_loopback_vlan_setting.txt # bcmcmd VLAN 30..137 建立 + 成對 pbmcdN <-> cdN+32
│ ├── traffic_loopback_start.txt # bcmcmd tx 100 length=512 VLantag=<vid>
── traffic_loopback_vlan_remove.txt # bcmcmd vlan remove <vid> pbm=...
# ↑ 三份為手動貼上用的原始指令表,
# blanton_traffic_linespeed.sh 是其 bash 版
── traffic_loopback_vlan_remove.txt # bcmcmd vlan remove <vid> pbm=...
# ↑ 三份為手動貼上用的原始指令表,
# blanton_traffic_linespeed.sh 是其 bash 版
│ ├── 100G_PRBS.txt # 100G PRBS 原始指令(port_prbs_monitor.sh 的來源)
│ ├── TR518.txt # 內建封包測試 tr 518Scenario=6 Profile=2)— 尚未腳本化
│ ├── fan_ctrl.txt # MAX31790 unbind/bind + fan-speed-control.sh 30
│ └── gen_channel_map.sh # settings/*.conf → docs/LTC2980_channel_map.csv(需 gawk
├── secret/ # 🚫 gitignored — DUT 密碼、per-unit BDF、COM 設定
│ ├── README.md # ✅ committed — 用途索引
@@ -148,8 +158,8 @@ Blanton_TTL_Script/
├── utils/ # 可 include 的取數片段(TTL,無狀態、只 wait+sendln
│ ├── pcie_bus.ttl # lspci -tvvv / -vv
│ ├── setup_pmon.ttl # show platform summary/fan/temperature/psustatus/voltage/
│ │ # current/ssdhealth/leak status/leak channels9 項)
│ ├── show_dmesg.ttl # date + dmesg grep error/fail/warning
│ │ # current/ssdhealth/leak status/leak channels + TPM11 項)
│ ├── show_dmesg.ttl # dmesg -T 合併 error 正則 + i2c 過濾 + dmesg -C 清空
│ ├── show_margin_status.ttl # source margin_status_all.sh(掃 9 顆 LTC2980
│ ├── kill_all_process.ttl # kill $(jobs -p)
│ ├── wait_init.ttl # 輪詢 show platform temperature 直到 pmon 起來
@@ -172,11 +182,17 @@ Blanton_TTL_Script/
├── blanton_pwr_data.sh # SWB VRM + PDB brick 的 Vin/Vout/Iout/Temp 表
├── blanton_traffic_linespeed.sh # SWB loopback 線速流量(bcmcmd unit 0/1
│ # init/clear/show/start/stop/ps/report/run
├── bmc_monitor.sh # BMC 監控:週期跑 free -m,寫輪替 log
├── bmc_monitor_ddr.sh # BMC DDR 壓力:memtester,逾時 3600s
│ # ↑ 兩支都要 chmod +x 才能跑(見 Key Constraints
├── mgmt_ping_monitor.sh # 10G/1G 管理網路 ping 監控(iproute2
│ # eth0→TARGET_10G、eth1→TARGET_1G,兩張卡同時持續 ping
├── port_prbs_monitor.sh # 100G uplink PRBS 測試(bcmcmd phy diag
│ # ⚠️ 已納入交付,但 A/B/C 的呼叫仍註解掉(bring-up 中)
├── usb_target.sh # 偵測 USB 裝置節點/掛載點(排除系統碟)
├── bmc_monitor.sh # BMC 監控:free -m + I2C 寫入/讀回圖樣測試
├── bmc_monitor_ddr.sh # ⚠️ 已停用:B/C 內呼叫處已註解,DDR 壓力改走 bgctl + bmc-manager
│ # ↑ 三支都要 chmod +x 才能跑(見 Key Constraints
└── LTC2980_Margin_Script/ # 電壓 margin 子專案(另有獨立 repo
├── margin.sh # margin_init/status/set/apply_profile/save
# # v2.7.0margin_save all / margin_save blanton
├── settings/*.conf # 一個 .conf = 一顆 LTC2980CB + SWB0/1 × CONN13~16
├── profiles/*.conf # 16-ch 批次組合(comboA/B、high3/5、low3/5、normal、off
├── script/*_all.sh # 跨 9 顆 LTC2980 的批次 status / apply / save
@@ -195,15 +211,20 @@ Blanton_TTL_Script/
### 1. `config.ttl` — 全域測試參數(TTL 變數)
```ini
strTestcase = "Margin" ; 測項名,進 log 檔名
project_name = "Blanton" ; 專案名,進 log 檔名
EN_Margin = 1 ; 1=跑 margin 掃描, 0=跳過
strTestcase = "ENV" ; 測項名(ENV/EMC/Margin...),進 log 檔名
EN_Margin = 0 ; 1=跑 margin 掃描, 0=跳過
EN_log = 1 ; 1=logopen 存檔, 0=不存
FAN_SPEED = 100 ; 傳給 fan-speed-control.sh 的風扇轉速 (%),預設全速
SWB_UNIT0 = 1 ; 這台有 switch unit 0(流量與 wait_init 都依此)
SWB_UNIT1 = 1 ; 這台有 switch unit 1
prompt_login = "sonic login:"
prompt_sonic = "admin@sonic:~$"
prompt_sonic_root = "root@sonic:~#"
```
> `SWB_UNIT0` / `SWB_UNIT1` 兩個都是 0 時,`wait_init.ttl` 直接 bypass、A/B/C 的流量段整段跳過。
Log 檔名規則:`<mdir>\Logs\<project_name>_<strTestcase>_<YYYYmmdd-HHMMSS>.log`
### 2. `settings/<board>.conf` — 一顆 LTC298016 channel
@@ -231,6 +252,31 @@ Channel 定址:`ch / 8` → `CHIPS[]` index`ch % 8` → LTC2977 PMBus PAGE
> 同一顆 LTC2980 的兩個位址**相差 2**7-bit,由規格書 8-bit 配對除以 2 而來),不是 1。
> 兩張板共用一條通道,所以同通道上的位址必須互斥 —— 上表兩組(`0x5C/0x5E` 與 `0x62/0x64`)不重疊。
**同一批 LTC2977 也掛在原生 i2c bus 上**2026-08-27 上機驗證):Blanton 的 SONiC 把 CB FPGA F3
的這四條 I2C 通道同時註冊成 kernel i2c adapter,因此除了 `cb_pmbus_*` → pcimem 這條路,
還可以用 `i2cset -f -y <bus>` 直接打到同一顆晶片:
| FPGA F3 通道 | Linux bus | CB FPGA 路徑 | 原生 i2c 路徑 |
|---|---|---|---|
| Ch6 | 13 | `cb_pmbus_write 6 <addr> …` | `i2cset -f -y 13 <addr> …` |
| Ch7 | 14 | `cb_pmbus_write 7 <addr> …` | `i2cset -f -y 14 <addr> …` |
| Ch8 | 15 | `cb_pmbus_write 8 <addr> …` | `i2cset -f -y 15 <addr> …` |
| Ch9 | 16 | `cb_pmbus_write 9 <addr> …` | `i2cset -f -y 16 <addr> …` |
`margin.sh``margin_save blanton` 走原生路徑(`MARGIN_BLANTON_STORE` 表),
其餘功能仍走 `TRANSPORT="swb"` 的 FPGA 路徑。兩條路到的是同一顆 LTC2977,
**不要同時用**(同一條匯流排上兩個 master 語意的存取)。
**全通道 net / 電壓對照:`docs/LTC2980_channel_map.csv`**9 個 `.conf` × 16ch = 144 列,
欄位 `Board,CONN,Ch,NetName,Vnom`)。查「某個 net 在哪片哪個 channel」比翻 9 個 `.conf` 快。
CSV 是產出物,不是正本 —— 改完 `.conf``./tools/gen_channel_map.sh` 重產,別手改 CSV。
兩件從表上看得出來、改 `.conf` 時要記得的事:
- **SWB0 與 SWB1 的 net / 電壓完全相同**,差別只在 `CHIPS[]` 的 bus`0:` vs `1:`)與 `CB_I2C_CH`
也就是說有兩份會各自漂移的重複資料 —— 改了 SWB0 沒改 SWB1,不會有任何東西擋你。
- **8 個 NC 通道**SWB0/1 CONN14 的 CH10/CH11/CH13、CONN15 的 CH10)在 CSV 裡以
`NC` / `-` 原樣保留,不是漏掉;濾掉它們會讓 page 編號對不上。
### 3. `profiles/*.conf` — 16-ch 批次 margin 組合
```bash
@@ -290,9 +336,9 @@ PROFILE_CHANNELS=( "0:high:8" "1:low:8" ... ) # "channel:operation:change_perc
| 模組 | 進入點 | 功能 |
|------|--------|------|
| **Script APre-test** | `1_Blanton_Script_A.ttl` V1.0.3 | root 登入 → `date -s` 對時 → `show boot`**等 pmon 起來**`wait_init.ttl`)→ BMC 版本 → source 六個 bash 工具 → 清/讀 7 組 PCIe AER → `lspci` → margin 全掃 → PMON 九項 → dmesg → 100G port 狀態 → **一輪流量基線**`show uptime` |
| **Script BStress** | `2_Blanton_Script_B.ttl` | `mlucas-avx2 -cpu 0:15` + 5 份 `stress_mem.py` + `stress_ssd.py` + `stress_usb.py`(**共 8 個 job**)→ 啟動兩支 BMC monitor → `jobs` → 100G port 狀態 → traffic `ps/init/clear/show/start`**兩個 unit**)→ `while 1` 每輪完整 PMON + margin |
| **Script CPost-test** | `3_Blanton_Script_C.ttl` | `kill $(jobs -p)` → PMON + dmesg → 7 組 AER → 停兩支 BMC monitor 並 `cat` 其 log`cat` mlucas / stress_ssd log → traffic `stop` + `report`(兩個 unit`show reboot-cause` / `show uptime`完成 messagebox |
| **Script APre-test** | `1_Blanton_Script_A.ttl` V1.0.7 | root 登入 → `date -s` 對時 → 清 job log → `hw-test-session start`**設定風扇轉速****等資料面就緒**`wait_init.ttl`)→ DUT 清單(boot/version/fwutil/syseeprom/ssdhealth/TPM/NVMe)→ BMC enroll + version/status → source bash 工具 → 清/讀 7 組 PCIe AER + `ras-mc-ctl --summary``lspci` → margin → PMON → dmesg → **10G/1G 管理網路 ping** → 100G uplink 設定與狀態 → **一輪流量基線**`show uptime` |
| **Script BStress** | `2_Blanton_Script_B.ttl` | `mlucas-avx2 -cpu 0:15` + `bgctl run``memtester` / SSD / USB 壓力 + BMC DDR`bmc-manager run memtester`)→ BMC USB net test(探測 cdc_ncm 介面,`systemd-run` 掛 4 小時 ping)→ 啟動 mgmt ping 監控 → `jobs` / `bgctl list` → traffic(依 `SWB_UNIT`)→ `while 1` 每輪完整 PMON + margin |
| **Script CPost-test** | `3_Blanton_Script_C.ttl` | `kill $(jobs -p)` + `bgctl stop --all` / `reset` + 停 BMC USB unit 與 mgmt ping → PMON + dmesg → 7 組 AER + `ras-mc-ctl` → BMC USB `journalctl``ip -s link show eth0/eth1``cat` mlucas log → traffic `stop` + `report` → NVMe 健康 `show reboot-cause` / `uptime`**job log 複製到 USB(帶時戳)**`hw-test-session finish` |
| **SWB loopback 流量** | `blanton_traffic_linespeed <sub> [-u 0\|1\|all]` | `init`VLAN 30..137,成對 `cdN`/`cdN+32`)→ `clear``start``tx 100 length=512`)→ `stop``report`per-pair TX/RX 交叉比對 PASS/FAIL);另有 `ps`link/speed)與 `run`(一條龍) |
| **FPGA 暫存器存取** | `cb_fpga <fn> <addr> [data]``fpga_rescan` | pcimem 打 `resource0` + BAR 相對 offset32-bit wordBDF 於 source 時自動偵測;`pmc/icb/swb0/swb1_fpga` 走 VSPI 窗口 |
| **CB I2C / PMBus** | `cb_i2c_init/scan/read/write``cb_pmbus_read/write/linear16` | FPGA F3 OpenCores master**讀用 Repeated START**SMBus/PMBus 裝置必需) |
@@ -300,7 +346,12 @@ PROFILE_CHANNELS=( "0:high:8" "1:low:8" ... ) # "channel:operation:change_perc
| **溫度** | `temp_cb` / `temp_icb` / `temp_all` | TMP7512-bit 左靠齊 ×0.0625)、TMP432(含 RANGE bit 的 64 偏移、remote1/2 |
| **電源資料** | `pwr_data` / `pwr_data_pdb` / `pwr_pdb_id` | SWB VRM 全軌 Vin/Vout/Iout/Temp 對齊表;MP29816 解析度**執行期從 `MFR_VOUT_SCALE_LOOP` 讀**,不寫死 |
| **電壓 Margin** | `margin_init/status/set/apply_profile/save` + `*_all.sh` | 9 顆 LTC2980CB CONN13、SWB0/1 CONN13~16)批次 high/low margin 與狀態掃描。v2.6.0 起 SWB 路徑真正接上共用 cb_i2c 後端,並有 `margin_swb_info/probe/scan/reset/sem` 除錯包裝 |
| **BMC 監控 / 壓力** | `bmc_monitor.sh` / `bmc_monitor_ddr.sh` `{start\|stop\|status\|fg\|tail\|clear}` | 由 host 端 `bmc-manager run` 週期性打 BMC,背景 detached、log 輪替。前者跑 `free -m`(逾時 30s),後者跑 `memtester`(逾時 3600s),各自獨立 log 與 pid |
| **Margin 寫 NVM** | `margin_save` / `margin_save all` / `margin_save blanton` | ⚠️ 永久寫入。`0x15` 送 PMBus **send-byte**(不帶 data byte)。`all` = 依 `settings/` 走 9 顆、跑完還原原本載入的板子;`blanton` = 本平台捷徑,18 顆全走原生 i2c bus(`MARGIN_BLANTON_STORE` 表),不需 `margin_init`。每筆之間 `MARGIN_STORE_SETTLE`0.5s),工具**不 poll busy bit** |
| **管理網路 ping** | `mgmt_ping_monitor.sh {start\|stop\|status\|fg\|tail\|clear\|summary}` | 一輪 = 10G leg + 1G leg,各自 `ip address/route replace``ping -I <if>` 打自己的對端;兩張卡都保持 up。每 leg 記 `ip -s link show` 與該卡的 TX/RX packet delta,另出一行可 grep 的 `RESULT` |
| **BMC 監控** | `bmc_monitor.sh {start\|stop\|status\|fg\|tail\|clear}` | 由 host 端 `bmc-manager run` 週期取樣 `free -m`,另加一組 **I2C 寫入/讀回圖樣測試**bus 1、addr 0x41、offset 0x00C0,寫 `55AA55AA` 讀回、再寫 `AA55AA55` 讀回)。B 啟動、C 停止並 `cat log/bmc_poll.log` |
| **BMC DDR 壓力** | `bmc_monitor_ddr.sh` | ⚠️ **已停用** —— B/C 內呼叫處已註解,改由 `bgctl run bmc-manager run memtester` 取代。檔案保留未刪 |
| **100G PRBS** | `port_prbs_monitor.sh {start [both\|a\|b]\|stop\|status\|report\|clear}` | 兩個 uplink 各一個 worker,可同時或單獨跑。setup`lpmode off``prbs set``prbsstat STArt`)→ 週期 `prbs get`(判定)+ `prbsstat Ber`(只記錄)→ teardown`STOp` + `prbs clear`)。輸出含 `PRBS OK!` 才算 PASS。**⚠️ A/B/C 的呼叫目前註解掉** |
| **USB 目標偵測** | `usb_target.sh [-o dev\|mnt\|both\|id] [-n] [-r sec]` | 找出插入的 USB 儲存裝置,排除 `/``/host` 的底層碟(本平台可能從 USB DOM 開機),多顆時拒絕猜;掛載時會驗證真的可寫 |
| **打包交付** | `publish/publish.sh [-v Ver] [-f] [-z] [-n]` | 版號讀 Script A 檔頭,複製成 `publish/Script_ABC_Blanton_<Ver>/`,排除 `*.log`,可另出 zip |
---
@@ -316,15 +367,20 @@ PROFILE_CHANNELS=( "0:high:8" "1:low:8" ... ) # "channel:operation:change_perc
Script A ── logopen Logs/Blanton_Margin_<ts>.log ──┐ (之後所有畫面都進這個檔)
│ │
├─ 登入 admin → sudo -i → root@sonic:~# │
├─ date -s "<host 時間>" ← 讓 DUT 時戳可對照
├─ date -s "<host 時間>" ← 讓 DUT 時戳可對照
├─ rm /host/hw-eval/jobs/* ← 清掉上一輪的 job log
├─ hw-test-session start / log / status
├─ 風扇:max31790 unbind→sleep→bind18-0020 與 25-0020 各一次)
│ → fan-speed-control.sh <FAN_SPEED> → 回答 y 確認 │
├─ source 6 個 bash 工具(函數進 shell
│ └ 此時 blanton_fpga_pcimem.sh 印出 [INFO] CB FPGA BDF = ...(自動偵測結果)
├─ AER 基線:7 組 EP+RP 的 correctable/nonfatal/fatal
├─ lspci -tvvv / -vv ← 拓樸 + link speed/width
├─ margin_status_all ← 9 顆 LTC2980 × 16ch 的 Vout 與偏差%
├─ show platform × 9 ← fan/temp/psu/voltage/current/ssdhealth/leak status/leak channels
├─ dmesg | grep error/fail/warning
├─ show interfaces status Ethernet513,Ethernet514 ← 100G uplink 狀態
├─ dmesg -T(合併 error 正則 + i2c 過濾)→ dmesg -C 清空 ← C 只會看到 soak 期間新產生的
├─ mgmt pingstart → pause 45 → stop → cat log ← 10G/1G 各一輪
├─ 100G uplinklpmode off → 兩埠 startup → config save → show interfaces status
├─ 一輪流量基線:ps → init → clear → show → start → stop → report
└─ show uptime
(最前面還有 wait_init.ttl:輪詢 show platform temperature 直到不再回報
@@ -336,8 +392,10 @@ Script A ── logopen Logs/Blanton_Margin_<ts>.log ──┐ (之後所有畫
Script B ── 同一個 log 檔續寫(logwrite 分隔線)
├─ 背景程序:mlucas-avx2(CPU) ×1, stress_mem ×5, stress_ssd, stress_usb → 共 8 個 job
stress_hhmd / stress_pcie 已於 2026-08-17 移除,見 Status 表 ScriptB #5/#7 "KC Remove"
├─ BMC: bmc_monitor.sh start + bmc_monitor_ddr.sh start
├─ jobs ← 確認全部起來了
├─ bgctl run: memtester 1G / qfx5252-stress-ssd / qfx5252-stress-usb
├─ BMC: bmc-manager run memtesterDDR)、cdc_ncm 介面探測 + systemd-run 4hr pingUSB
├─ mgmt_ping_monitor.sh start
├─ jobs / bgctl list ← 確認全部起來了
├─ traffic: ps → init → clear → (pause 15) → show → start unit 0 與 1VLAN 30..137 loopback
└─ while 1: setup_pmon9 項)→ margin_status_all soak 期間持續取樣)
@@ -348,10 +406,15 @@ Script C
├─ kill $(jobs -p) ← 收掉所有背景壓力
├─ show platform × 9 + dmesg ← 收工快照
├─ AER 再抓一次 ← 與 Script A 基線相減 = 本輪新增錯誤
├─ BMC: 兩支 monitor stop → cat log/bmc_poll.log + log/bmc_ddr.log
├─ cat mlucas_amm_log / log.stress_ssd ← 壓力程式自身的 pass/fail
├─ bgctl stop --all / reset、systemctl stop BMC USB unit、mgmt ping stop
├─ journalctl BMC USB unit、ip -s link show eth0/eth1、ras-mc-ctl --summary
├─ cat mlucas_amm_log ← CPU 壓力程式自身的 pass/fail
├─ nvme smart-log / smartctl -x /dev/nvme0
├─ traffic: stop → report ← per-pair TX/RX 交叉比對,PASS/FAIL/NAexit 2 = 有 FAIL/NA
─ show reboot-cause / show uptime ← soak 期間有沒有意外重開,一眼看得到
─ show reboot-cause / show uptime ← soak 期間有沒有意外重開,一眼看得到
├─ usb_target.sh -o dev → sync → mount <dev> /mnt/usb
├─ cp bmc_poll.log / mgmt_ping.log → jobs/,再 cp -r jobs/ /mnt/usb/jobs-<YYYYmmdd>_<HHMM>
└─ hw-test-session finish
[messagebox "GOOD JOB! Test Case DONE"] → 人工檢查 log
```
@@ -389,6 +452,8 @@ pwr_data
只吃 Sr;只有純暫存器裝置兩種都行。
7. **`margin_save``STORE_USER_ALL`)會永久寫進 NVM**,斷電保留。測試流程中**絕不執行**,
以免覆蓋客戶出廠 config。`margin.sh` 的 auto-enable 只改 RAM`ON_OFF_CONFIG`),斷電復原。
`margin_save all` / `margin_save blanton`v2.7.0)一次對 9 顆 LTC298018 個 LTC2977
下同一道命令,風險同性質但範圍是全機;命令本身是 PMBus send-byte`0x15`,不帶 data byte)。
8. **操作硬體電壓有損壞 DUT 的風險**:套 profile 前先確認該軌 OV/UV limit 與 servo DAC 注入電阻
有 populate(本板部分 margin 電阻標 PROTO,沒 populate 的軌電壓不會動)。
9. **`ARB_LOST` 不等於裝置存在**`vi2c_scan` 只看 `RX_ACK`,仲裁失敗會被誤報成「有裝置」。
@@ -409,12 +474,21 @@ pwr_data
⚠️ 但 `.gitattributes` **不會回頭重寫既有工作區檔案** —— 早於它 clone 出來的目錄,
`git rm --cached -r . && git reset --hard` 或刪掉目錄重新 checkout 才會生效,
而且 `git status` 乾淨**不代表**磁碟上是 LF`--renormalize` 只在進 index 的路上轉換)。
13. **`bmc_monitor*.sh` 需要執行權**:它們是本專案第一組「被執行」而非被 `source` 的腳本,
13. **`bmc_monitor*.sh` / `mgmt_ping_monitor.sh` 需要執行權**:它們是本專案「被執行」而非被 `source` 的腳本,
`do_start` 內部是 `setsid "$SCRIPT_PATH" __daemon`**即使用 `bash x.sh start` 呼叫
也仍需 exec bit**。git 記錄為 `100644`,且 Windows → 隨身碟 → DUT 的路徑無法攜帶權限位元,
所以佈署後要先 `chmod +x ~/Blanton_Script/bmc_monitor*.sh`,否則 Script B 的 start 會
`Permission denied`、Script C 的 `cat` 撲空(BMC 段全空但不會報錯)。
14. **客戶 NDA**`docs/` 內含客戶 FPGA 規格與 schematic 衍生資訊,**本 repo 只推 NAS + Gitea
14. **TTL 的 `timeout` 是全域的,設了就要還原**`timeout = N` 之後**每一個** `wait` 都被限制在 N 秒
逾時的 `wait` 會直接返回而沒吃到 prompt,於是下一個指令在前一個還沒跑完就送出,之後整輪錯開一拍。
`wait_init.ttl` 設 60 後還原 0Script B 的 cdc_ncm 探測設 15,在 `:skip_ping`(兩條路徑的匯流點)還原。
加新的 `timeout` 時務必配一個 `timeout = 0`
15. **兩張管理網卡同網段時會 ARP flux**`mgmt_ping_monitor.sh` 讓 eth0/eth1 同時 up,若兩者在同一個
/24(本專案預設 `192.168.1.99` / `.101`),對端 ARP 可能由任一張卡回應,回包落在沒送封包的那張,
量到的數字不能歸屬。`ARP_STRICT=1` 會設 `arp_ignore=1` / `arp_announce=2` 擋掉;
另外每個 leg 的 `RESULT` 行帶 `nic_tx` / `nic_rx`(該卡自己的 packet delta),
**ping 成功但 `nic_tx` 接近 0 就代表封包從另一張卡出去了**。兩張卡在不同網段時可設 0。
16. **客戶 NDA**`docs/` 內含客戶 FPGA 規格與 schematic 衍生資訊,**本 repo 只推 NAS + Gitea
絕不推 GitHub**。
---
@@ -428,7 +502,7 @@ pwr_data
| **客戶機密文件** | `docs/` 內 FPGA 規格(docx/xlsx/md)屬客戶 NDA 範圍;schematic 原始檔(`*.DSN`)一律 gitignore,不進 repo |
| **Remote 限制** | 只推私有 remote`nas``git.et-wen.com:51222`+ `gitea``etgit.et-wen.com`)。**不設 GitHub `origin`** |
| **測試 log** | `logs/*.log` 含 DUT 序號、MAC、韌體版本等可識別資訊 → gitignore,只保留目錄結構(`.gitkeep` |
| **破壞性硬體指令** | `margin_save` / `STORE_USER_ALL` 為永久 NVM 寫入,工具預設不呼叫;批次腳本 `margin_save_all.sh` 在 README 與腳本頂端都有警語 |
| **破壞性硬體指令** | `margin_save` / `margin_save all` / `margin_save blanton` / `STORE_USER_ALL` 為永久 NVM 寫入,工具預設不呼叫;`margin_save all`批次腳本 `margin_save_all.sh` 在 README 與程式頂端都有警語 |
| **root 權限** | 全流程在 `root@sonic:~#` 下執行(pcimem / i2c 需要);腳本不留任何 `chmod 777` 或放寬權限的動作 |
| **For_AI/** | AI 協作素材(截圖、草稿筆記)整個資料夾 gitignore,避免不小心把客戶圖面帶進 repo |
@@ -554,16 +628,17 @@ margin status、全部 VRM 軌的 Vin/Vout/Iout/Temp、CB + ICB 溫感讀值,
**目標:** 壓力程序起得來、起不來看得出來,soak 迴圈有時戳、可控週期。
**包含:**
- [ ] 部署 `~/hammer/tools/` 前置檢查:新增 `utils/check_hammer_tools.ttl`
`ls -l ~/hammer/tools/` + `ls ~/hammer/tools/amd/mlucas-avx2`),缺檔就別往下跑
- [ ] 前置檢查:`~/hammer/tools/amd/mlucas-avx2` 仍是唯一還依賴 hammer 的項目,缺檔就別往下跑
- [ ] 壓力程序 log 加時戳,避免多輪覆蓋:
`mlucas_amm_log``~/logs/mlucas_amm_log_$(date +%Y%m%d-%H%M%S).log`Status 表 ScriptB #2 的既定寫法)
- [ ] `jobs` 之後加驗證:期望 **8** 個背景 job1 mlucas + 5 mem + ssd + usb),
數量不符時 `messagebox` 提示
`stress_hhmd` / `stress_pcie` 已於 2026-08-17 由 KC 移除,原本的 10 改為 8
- [x] 壓力層改用平台的 `bgctl``memtester` / `qfx5252-stress-ssd` / `qfx5252-stress-usb`),
取代 `~/hammer/tools/stress_*.py`Script B 以 `jobs` + `bgctl list` 兩者並列確認
- [ ] `bgctl list` 之後加數量驗證,不符時 `messagebox` 提示(目前只印出來給人看
**← V1.0.10 起 `bgctl list` + `jobs` 改為每輪都印,中途死掉的 job 看得到,但仍要靠人眼比對**
- [ ] soak 迴圈改為「每輪先 `date`、跑完 `pause <間隔>`」,
`EN_LoopInterval`(預設 600 秒 = 10 min)放進 `config.ttl`,兌現 Status 表的「Get data every 10mins」
**目前 `while 1` 仍無 `pause`,實際間隔取決於指令執行時間,不是 10 分鐘**
**V1.0.10 已補上 `pause 60`(迴圈不再全速空轉),但 60 秒是寫死的、也還沒有每輪 `date`
距 Status 表的 10 分鐘仍有差距,故本項未結**
- [x] soak 每輪補 `show platform psustatus` / `voltage` / `current`
—— 迴圈已改為 `include "utils/setup_pmon.ttl"`,九項全收
- [ ]`; ========== CLEAR EVENT ==========`Status 表 ScriptB #1Owner: Alan
@@ -621,9 +696,13 @@ margin status、全部 VRM 軌的 Vin/Vout/Iout/Temp、CB + ICB 溫感讀值,
UE > 0 或 fatal/nonfatal > 0 直接標 FAIL
- [ ] Margin 表解析:把 9 顆 × 16ch 的 Vout / 偏差% 收成 CSV,標出超出 ±(profile 設定值 + 容差) 的軌
- [ ] 溫度 / 風扇 / 電源趨勢:soak 期間逐輪取樣 → CSV(可直接畫圖)
- [ ] Stress 判定:`mlucas_amm_log``log.stress_ssd` 的錯誤關鍵字掃描
- [ ] Stress 判定:`mlucas_amm_log``bgctl` 各工作的 log
`/host/hw-eval/.../qfx5252-stress-{ssd,usb}.log`)的錯誤關鍵字掃描
- [ ] Traffic 判定:**直接複用 `tools/bcm_mibpair_report_V1.1.0.py`**(它已能吃 console log 的
`show c` 段落並輸出 PASS/FAIL),不要重寫一套解析
- [ ] 管理網路判定:抓 `mgmt_ping_monitor.sh``RESULT` 行(已是單行可 grep 格式),
並檢查 `nic_tx`/`nic_rx` 不為 0 —— 否則 PASS 也不能歸屬到那張卡
- [ ] NVMe / EDAC`nvme smart-log``smartctl -x``ras-mc-ctl --summary` 的前後差值
- [ ] 產出 `Logs/<run>_summary.md`:一頁 PASS/FAIL + 每個子系統一行結論
**驗收條件:** 拿一份既有 log(如 `Blanton_Margin_20260814-164122.log`)跑解析器,
@@ -676,8 +755,36 @@ margin status、全部 VRM 軌的 Vin/Vout/Iout/Temp、CB + ICB 溫感讀值,
文件已加過期警告但內文未動。第 1~4 章不受影響。
- **`wait_init.ttl` 沒有逾時出口**:若 pmon 一直起不來,`result` 既非 1 也非 2,迴圈會無限重試,
Script A 既不往下走也不報錯。可加最大重試次數 + `messagebox` 提示。
- **`bmc_monitor*.sh` 的執行權**:目前需人工 `chmod +x`。可在 Script B 起它們之前補一行
`chmod +x`(比照既有的 `chmod +x ~/hammer/tools/amd/mlucas-avx2`),或在 publish 時處理。
- **`bmc_monitor_ddr.sh` 的壓力強度**`memtester 500M 1``INTERVAL_SEC=5` 等於整個 soak
近乎不間斷壓 BMC 記憶體,而 BMC 同時還要服務 host 的 `show platform` 查詢
若 PMON 取數變慢,先調大 `INTERVAL_SEC` 或縮小 memtester 的量。
- **腳本的執行權**`mgmt_ping_monitor.sh``bmc_monitor.sh`(以及已停用的 `bmc_monitor_ddr.sh`)需人工 `chmod +x`
可在 Script B 起它之前補一行(比照既有的 `chmod +x ~/hammer/tools/amd/mlucas-avx2`),或在 publish 時處理。
- **`bmc_monitor_ddr.sh` 已停用**:B/C 內的呼叫已註解,DDR 壓力改由 `bgctl run bmc-manager run memtester`
取代。檔案還留在 repo,待確認不再需要後刪除,或在檔頭標註停用
`bmc_monitor.sh` 已於 2026-08-25 重新啟用,並加上 I2C 寫入/讀回圖樣測試。)
- **`rm /host/hw-eval/jobs/*` 清不掉子目錄**:Script A 用的是不帶 `-r``rm`,但 Script C
`cp -r` 那個目錄,代表裡面可能有子目錄。清不乾淨的話,上一輪的 job log 會被當成這輪的一起
複製到 USB。要確認 `bgctl` 的產出結構,必要時改成 `rm -rf .../jobs/*`
- **100G PRBS 尚未接進 A/B/C**`port_prbs_monitor.sh` 已隨包交付、可手動執行,但三支腳本裡的呼叫
都還註解著(bring-up 中)。要啟用時記得 PRBS 期間 link 會顯示 down`stop` 之後才會回來。
- **Script C 的 USB 掛載走 `mount` 而非 `usb_target.sh -o mnt`**:目前是 `-o dev` 取節點再自己
`mount <dev> /mnt/usb`,繞過了工具內建的 `mkdir -p`、exfat/ntfs `modprobe` 與可寫性驗證。
掛載點不存在、檔案系統模組沒載、或媒體唯讀時只會失敗一行,job log 就沒帶出來。
- **風扇確認提示是無限等待**`wait "Set all configured fan channels to"` 之後才 `sendln "y"`
若該版 `fan-speed-control.sh` 不再詢問(或字串改了),這個 `wait` 會永遠等下去,Script A 停在那裡
且不報錯。加個 `timeout` + 分支會比較保險(記得還原 `timeout = 0`)。
- **TR518 內建封包測試尚未腳本化**:`tools/TR518.txt` 記下了 `tr 518 Scenario=6 Profile=2 Count=260
PktSize=324 PortList=1-259,270-565` 這組指令(含 `port all lb=mac` + `l2 learn off` 前置),
目前只能手貼。它與 `blanton_traffic_linespeed` 是兩條互斥的打流路線(都會動 loopback 與 L2 學習),
要腳本化的話得先決定兩者怎麼共存,不能同時跑。
- **SWB 的 LTC2977 有兩條路可以到,目前工具兩條都在用**:`cb_pmbus_*`CB FPGA F3 → pcimem
是 `margin_init` / `margin_status` / `margin_set` 走的路;原生 i2c busCh+7 = bus 13/14/15/16
是 `margin_save blanton` 走的路。後者短很多 —— 不需要通道初始化、不需要 pcimem,
單一 `i2cset` 就到。若原生 bus 在這平台夠穩,整個 `TRANSPORT="swb"` 層其實可以退成
「bus 號不同的 i2c 路徑」,`_swb_*`、probe、semaphore、prescale 那整套都能拿掉。
要動之前得先確認:讀取(`margin_status` 的 LINEAR16 word read)是否也同樣可靠,
以及 kernel driver 有沒有搶著存取(`margin_save blanton` 帶 `-f` 就是為了繞過這件事)。
- **`settings/` 裡 SWB0 與 SWB1 是兩份相同的資料**:8 個 `.conf` 中,SWB1 的四個除了 `CHIPS[]` 的
bus 與 `CB_I2C_CH` 之外,net name 與 `VNOM` 與 SWB0 逐欄相同(見 `docs/LTC2980_channel_map.csv`)。
改一邊忘了改另一邊不會有任何警告。可考慮改成「共用 net 定義 + 各自的 transport 覆寫」兩層檔案。
- **Script C 沒有倒出 `mgmt_ping.log`**C 會 `stop` 管理網路監控並印 `ip -s link show`
但沒有 `cat ~/Blanton_Script/log/mgmt_ping.log`,所以整個 soak 期間累積的 per-leg `RESULT`
留在 DUT 上、沒進 master log。Script A 那輪基線有 `cat`。
+81 -16
View File
@@ -9,9 +9,16 @@ DUT 端的取數能力由一組 **bash 工具庫** 提供(CB FPGA `pcimem` →
完整架構見 [ARCHITECTURE.md](ARCHITECTURE.md)。
**目前狀態(2026-08-21**Script A/B/C**V1.0.3**,涵蓋 PCIe AER / DDR EDAC / PMON /
電壓 margin9 顆 LTC2980/ SWB loopback 線速流量(兩個 switch unit/ BMC 監控與 DDR 壓力。
SWB 的 I2C 通道與位址已於 2026-08-19 上機驗證。交付用 `./publish/publish.sh` 打包
**目前狀態(2026-08-27**Script A 為 **V1.0.10**(A 集中記錄 A/B/C 三支的變更)。涵蓋
PCIe AER / DDR EDAC / PMON+TPM / 電壓 margin9 顆 LTC2980/ SWB loopback 線速流量 /
10G+1G 管理網路 ping / BMC DDR 與 USB / NVMe 健康,壓力層走平台的 `bgctl`
switch unit 數量由 `config.ttl``SWB_UNIT0` / `SWB_UNIT1` 宣告,風扇轉速由 `FAN_SPEED` 指定
(預設 100 = 全速;soak 是熱測試,散熱不該是變因)。
測試 job log 收在 DUT 的 `/host/hw-eval/jobs/`Script C 結束時掛好 USB 再帶時戳複製到 `/mnt/usb/`
100G PRBS`port_prbs_monitor.sh`)已隨包交付,但 **A/B/C 的呼叫仍註解掉**,還在 bring-up。
Script B 的 soak 迴圈每輪印 `bgctl list` + `jobs`,並 `pause 60`V1.0.10 之前完全沒有 pause)。
最後一次發布是 **V1.0.9**;SWB 的 I2C 通道與位址已於 2026-08-19 上機驗證。
交付用 `./publish/publish.sh` 打包。
## 技術棧
@@ -24,6 +31,9 @@ SWB 的 I2C 通道與位址已於 2026-08-19 上機驗證。交付用 `./publish
- **資料面流量**`bcmcmd`Broadcom drivshellSWB loopbackVLAN 30..137 成對 `cdN`/`cdN+32`
`tx 100 length=512`;判定讀 ASIC `MIB_TPOK`/`MIB_RPOK` 做 per-pair 雙向交叉比對
- **BMC**host 端 `bmc-manager run "<cmd>"`,不需對 BMC 開 SSH(繞開其 busybox 工具鏈)
- **背景工作**:平台的 `bgctl run/list/stop --all/reset`;長時間 ping 用 `systemd-run --unit=` 掛 transient unit
- **管理網路**iproute2`ip link set` / `address replace` / `route replace ... metric`+ `ping -I`
- **風扇**max31790 driver rebind + `fan-speed-control.sh <%>`(轉速讀 `config.ttl``FAN_SPEED`
- **判定資料源**PCIe AER sysfs`aer_dev_*`)、EDAC sysfs`dimm_{ce,ue}_count`)、SONiC `show platform *`
- **無編譯步驟**:TTL 與 bash 都直譯執行
@@ -54,6 +64,14 @@ temp_all # CB + ICB 溫感
pwr_data # SWB VRM 全軌
margin_init Blanton_CB_CONN13 && margin_status
# ── 電壓 margin 寫 NVM(⚠️ 永久,測試中絕不執行)─────
margin_save # 只寫當前 margin_init 載入的那顆(2 個 LTC2977
margin_save all # settings/ 全部 9 顆;跑完自動切回原本載入的那顆
margin_save blanton # Blanton SONiC 捷徑:18 顆全走原生 i2c,不用 margin_init
# 0x15 是 send-byte,不帶 data bytei2cset -f -y <bus> <addr> 0x15
# bus = FPGA F3 通道 + 7 → Ch6=13, Ch7=14, Ch8=15, Ch9=16CB 本來就是 bus 4
# ⚠️ 沒裝滿的機器:改 margin.sh 頂端的 MARGIN_BLANTON_STORE,把沒有的板子註解掉
# ── SWB loopback 流量(bcmcmdunit 0/1)─────────────
# 不帶 -u 就是預設 TL_UNITS="0 1",兩塊 switch board 都做(TTL 三支腳本即如此)
blanton_traffic_linespeed ps # 對照線的 link/speed
@@ -70,13 +88,38 @@ blanton_traffic_linespeed report -f <log> # 離線重解一份存下來的 cons
# 離線版報表(開發機上跑,選項更多)
python3 tools/bcm_mibpair_report_V1.1.0.py <drivshell.log> -a
# ── BMC 監控 / 壓力(需先 chmod +x!)────────────────
chmod +x ~/Blanton_Script/bmc_monitor*.sh # 佈署後必做一次
~/Blanton_Script/bmc_monitor.sh start # free -m 取樣,log/bmc_poll.log
~/Blanton_Script/bmc_monitor_ddr.sh start # memtesterlog/bmc_ddr.log
~/Blanton_Script/bmc_monitor.sh status # pid / log 路徑 / log 大小
~/Blanton_Script/bmc_monitor.sh stop
# ⚠️ tail 子命令是 tail -f,會卡住不返回;要倒 log 請直接 cat log/*.log
# ── 管理網路 ping 監控(需先 chmod +x!)──────────────
chmod +x ~/Blanton_Script/*.sh # 佈署後必做一次
~/Blanton_Script/mgmt_ping_monitor.sh fg # 前景試一輪(約 30 秒),確認 IP/對端/網卡
~/Blanton_Script/mgmt_ping_monitor.sh start # 背景,log/mgmt_ping.logstart 會先清空)
~/Blanton_Script/mgmt_ping_monitor.sh status
~/Blanton_Script/mgmt_ping_monitor.sh stop
~/Blanton_Script/mgmt_ping_monitor.sh summary # 只看每段的 RESULT 行
# ⚠️ tail 子命令是 tail -f,會卡住不返回;TTL 裡要倒 log 請用 cat
# bmc_monitor.sh start/stop # free -m + BMC I2C 寫入/讀回圖樣測試,log/bmc_poll.log
# ⚠️ bmc_monitor_ddr.sh 已停用(DDR 壓力改走 bgctl + bmc-manager),檔案仍在
# ── 100G PRBS(目前 TTL 未啟用,手動測用)────────────
bash ~/Blanton_Script/port_prbs_monitor.sh start # 兩個 port 同時
bash ~/Blanton_Script/port_prbs_monitor.sh start a # 只跑 Ethernet513
bash ~/Blanton_Script/port_prbs_monitor.sh start b # 只跑 Ethernet514
bash ~/Blanton_Script/port_prbs_monitor.sh report # PASS/FAIL 統計(沒跑的顯示 SKIP
bash ~/Blanton_Script/port_prbs_monitor.sh stop # 收尾會跑 prbsstat STOp + prbs clear
# ⚠️ PRBS 打起來時 link 會顯示 down,這是正常的 —— 所以 start 刻意不印 port status
# ── USB 目標偵測 ────────────────────────────────────
bash ~/Blanton_Script/usb_target.sh -o dev # /dev/sdX1
bash ~/Blanton_Script/usb_target.sh -o mnt # 掛好並回傳掛載點
# ── 背景壓力(平台工具)──────────────────────────────
bgctl run /usr/sbin/memtester 1G 100
bgctl list
bgctl stop --all ; bgctl reset --yes
# ── 通道對照表(開發機上跑,改完 settings/*.conf 一定要重產)──
./tools/gen_channel_map.sh # → docs/LTC2980_channel_map.csv144 列)
./tools/gen_channel_map.sh - # 只印到 stdout
# ⚠️ 需要 gawk(用到 3 參數的 match());mawk 會直接擋下並提示
# ── 各工具的 help ───────────────────────────────────
blanton_fpga_help ; blanton_cb_i2c_help ; vi2c_help
@@ -144,6 +187,11 @@ Tera Term 連 COM port (115200-8-N-1)
- ⚠️ **I2C 讀一律 Repeated START**(不是 STOP + START),否則 PMBus/SMBus 裝置(VRM、EFUSE)不理。
- ⚠️ **`margin_save` / `STORE_USER_ALL` 會永久寫 NVM**,斷電保留 → **測試流程中絕不執行**
以免覆蓋客戶出廠 config。`margin.sh` 的 auto-enable`ON_OFF_CONFIG=0x1a`)只改 RAM,斷電復原。
v2.7.0 新增的 **`margin_save all` / `margin_save blanton` 一次寫 18 個 LTC2977**
同一個警告放大 9 倍;另外 `0x15` 現在送 send-byte,不再帶多餘的 `0x00` data byte。
- ⚠️ **SWB 的 LTC2977 有兩條路可以到**`cb_pmbus_*`CB FPGA F3 Ch6/7/8/9 → pcimem)與
原生 i2c bus`i2cset -f -y 13/14/15/16`Ch+7 = bus)。`margin_save blanton` 走後者,
其餘功能走前者。同一顆晶片,**不要同時用兩條路**。
- ⚠️ **改電壓會弄壞 DUT**:套 profile 前確認該軌 OV/UV limit,且 servo DAC 注入電阻有 populate
(本板部分 margin 電阻標 PROTO,沒 populate 的軌電壓不會動,不是腳本壞掉)。
- ⚠️ **`ARB_LOST` ≠ 有裝置**`vi2c_scan` 只看 `RX_ACK`,仲裁失敗會被誤報成裝置存在。
@@ -166,13 +214,27 @@ Tera Term 連 COM port (115200-8-N-1)
`.gitattributes` 已擋住 git 這條路徑,但**不會回頭重寫既有工作區檔案**:早於它 clone 的目錄要
`git rm --cached -r . && git reset --hard` 才會生效,而且 **`git status` 乾淨不代表磁碟上是 LF**
`--renormalize` 只在進 index 的路上轉換)。驗證要看 byte:`grep -rlU $'\r' <dir>`
- ⚠️ **`bmc_monitor*.sh` 需要 exec bit**,且**不是改用 `bash x.sh start` 就能繞過** ——
- ⚠️ **TTL 的 `timeout` 是全域的**`timeout = N` 之後**每一個** `wait` 都被限制在 N 秒。逾時的 `wait`
直接返回而沒吃到 prompt,下一個指令就在前一個還沒跑完時送出,之後整輪錯開一拍。**設了一定要
配一個 `timeout = 0` 還原**`wait_init.ttl` 60→0Script B 的 cdc_ncm 探測 15→0 在 `:skip_ping`
- ⚠️ **兩張管理網卡同網段會 ARP flux**`mgmt_ping_monitor.sh` 讓 eth0/eth1 同時 up,同一個 /24 下
對端 ARP 可能由任一張卡回應,回包落在沒送封包的那張。`ARP_STRICT=1` 擋掉;判讀時看 `RESULT` 行的
`nic_tx`/`nic_rx`(該卡自己的 packet delta)—— **ping 成功但 `nic_tx` 接近 0 = 封包從另一張卡出去的**
- ⚠️ **PRBS 跑起來時 `show interfaces status` 會顯示 DOWN**,那是 link 離開正常運作模式,不是故障。
`port_prbs_monitor.sh``start` 因此刻意不印 port status(避免測試員誤判),只在 `prbs clear` 之後的報表裡印
- ⚠️ **`ip -s link` 的欄位順序是 `bytes packets errors ...`**,要 packet 數取的是**第 2 欄**不是第 1 欄
- ⚠️ **`mgmt_ping_monitor.sh` / `bmc_monitor*.sh` 需要 exec bit**,且**不是改用 `bash x.sh start` 就能繞過** ——
`do_start` 內部是 `setsid "$SCRIPT_PATH" __daemon`,直接執行該路徑。git 記錄是 `100644`
Windows → 隨身碟 → DUT 也帶不動權限位元,所以佈署後要 `chmod +x ~/Blanton_Script/bmc_monitor*.sh`
沒做的話 Script B 的 start 會 Permission denied、Script C `cat` 撲空,而且**不會報錯**
只會得到一份 BMC 段全空的 log
- ⚠️ **`bmc_monitor*.sh``tail` 子命令是 `tail -n 50 -f`,永遠不返回**。TTL 裡絕對不要呼叫它
沒做的話 start 會 Permission denied、後面`cat` 撲空,而且**不會報錯**只會得到一份全空的 log
- ⚠️ **這幾支 monitor 的 `tail` 子命令是 `tail -n 50 -f`,永遠不返回**。TTL 裡絕對不要呼叫它
Script C 曾因此卡死,`18ce7a1` 改成直接 `cat log/*.log`
- ⚠️ **背景 monitor 的 `start` 只等 1 秒就返回**,不是跑完一輪。Script A 要做一輪基線的話,
`start``stop` 中間必須 `pause`(目前是 45 秒,一輪約 30 秒)—— 少了它 log 只會有 START banner
- ⚠️ **風扇那段設完會 `wait "Set all configured fan channels to"` 等確認提示再送 `y`**。該提示沒出現
(版本改了、字串換了)就會永遠卡住且不報錯 —— 現場看起來像 Script A 當掉
- ⚠️ **`show_dmesg.ttl` 結尾會 `dmesg -C` 清空 kernel ring buffer**。這是刻意的(讓 Script C 只看到
soak 期間新產生的訊息),但代表事後在 DUT 上 `dmesg` 撈不到舊訊息 —— 內容只存在 master log 裡
- ⚠️ **DUT 的 awk 對 `%d` 會夾到 INT32**(2147483647)。線速流量的計數器是 ~10¹² 量級,
所以 awk 裡輸出大數一律用 `%.0f`double 域,精確到 2⁵³)。`blanton_traffic_linespeed.sh`
V0.4.1 已修;寫新的 awk 報表時要記得同一件事
@@ -180,10 +242,13 @@ Tera Term 連 COM port (115200-8-N-1)
## 資料夾說明
- `src/Script_ABC_Blanton/` — 可執行內容(Tera Term 工作目錄;內含要 scp 到 DUT 的 `Blanton_Script/`
- `docs/` — 客戶 FPGA 規格、AER/EDAC 檢查依據、逐項 Status 表(最新為 `*_20260817.xlsx`
- `docs/` — 客戶 FPGA 規格、AER/EDAC 檢查依據、逐項 Status 表(最新為 `*_20260817.xlsx`
`LTC2980_channel_map.csv`9 個 `settings/*.conf` 合併成 Board,CONN,Ch,NetName,Vnom 144 列,
查 net 對應哪片哪個 channel 用;**改 `.conf` 後要重跑產生**
- `tools/` — Host / 離線端:`bcm_mibpair_report_V1.1.0.py`drivshell log → per-pair 報表)、
`traffic_loopback_*.txt`bcmcmd 原始指令表,`blanton_traffic_linespeed.sh` 是其 bash 版,
**改 loopback 線路時兩邊要同步**
**改 loopback 線路時兩邊要同步**`100G_PRBS.txt` / `TR518.txt` / `fan_ctrl.txt`
(上機手貼用的原始指令;PRBS 已有 `.sh`TR518 尚未腳本化)
- `publish/``publish.sh`(打包工具,**進 git**+ `Script_ABC_Blanton_<Ver>/` 產出(**gitignored**
- `secret/` — 🚫 gitignored(除 `README.md``*.example`):DUT 帳密、per-unit BDF 覆寫、COM 設定
- `For_AI/` — 🚫 gitignored:AI 協作素材(波形截圖、草稿筆記)
+145
View File
@@ -0,0 +1,145 @@
Board,CONN,Ch,NetName,Vnom
CB,CONN13,CH0,V5P0_ALW,5.0
CB,CONN13,CH1,V3P3_ALW,3.3
CB,CONN13,CH2,PWR_VDD_MISC_ALW,0.75
CB,CONN13,CH3,V1P8_ALW,1.8
CB,CONN13,CH4,PWR_APU_VDDIO_SUS,1.1
CB,CONN13,CH5,PWR_VDD_MISC_RUN,0.75
CB,CONN13,CH6,PWR_APU_VDD_MEM_RUN,0.78
CB,CONN13,CH7,V0P8_PHY2_DVDD,0.8
CB,CONN13,CH8,V0P8_AVDD,0.8
CB,CONN13,CH9,V0P8_PHY,0.8
CB,CONN13,CH10,V1P8_FPGA,1.8
CB,CONN13,CH11,V2P5_FPGA,2.5
CB,CONN13,CH12,V1P1_FPGA,1.1
CB,CONN13,CH13,P5V_STBY,5.0
CB,CONN13,CH14,P3V3_STBY,3.3
CB,CONN13,CH15,V0P8_PHY2_AVDD,0.8
SWB0,CONN13,CH0,P0V75_DVDD_47,0.75
SWB0,CONN13,CH1,P0V75_AVDD_47,0.75
SWB0,CONN13,CH2,P1V5_AVDD_47,1.5
SWB0,CONN13,CH3,P0V9_AVDD_47,0.9
SWB0,CONN13,CH4,P1V8_1,1.8
SWB0,CONN13,CH5,PVDD3_3V_SYNTH_LDO_A3,3.3
SWB0,CONN13,CH6,PVDD3_3V_SYNTH_LDO_B1,3.3
SWB0,CONN13,CH7,P1V8_3,1.8
SWB0,CONN13,CH8,P1V8_DVDDIO_3,1.8
SWB0,CONN13,CH9,P0V75_DVDD_39,0.75
SWB0,CONN13,CH10,P0V75_AVDD_39,0.75
SWB0,CONN13,CH11,P1V5_AVDD_39,1.5
SWB0,CONN13,CH12,P0V9_AVDD_39,0.9
SWB0,CONN13,CH13,PVDD1V5_2,1.5
SWB0,CONN13,CH14,PVDD1V8_DUT,1.8
SWB0,CONN13,CH15,PVDD_MDIO,1.2
SWB0,CONN14,CH0,P0V75_AVDD_7,0.75
SWB0,CONN14,CH1,P1V5_AVDD_7,1.5
SWB0,CONN14,CH2,P0V9_AVDD_7,0.9
SWB0,CONN14,CH3,P1V8_5,1.8
SWB0,CONN14,CH4,P1V8_2,1.8
SWB0,CONN14,CH5,PVDD3_3V_SYNTH_LDO_B2,3.3
SWB0,CONN14,CH6,PVDD3_3V_SYNTH_LDO_A2,3.3
SWB0,CONN14,CH7,PVDD3_3V_SYNTH_LDO_A1,3.3
SWB0,CONN14,CH8,PVDD1V5_TSC,1.5
SWB0,CONN14,CH9,PVDD1V5_ANLG,1.5
SWB0,CONN14,CH10,NC,-
SWB0,CONN14,CH11,NC,-
SWB0,CONN14,CH12,PVDD1V5_1,1.5
SWB0,CONN14,CH13,NC,-
SWB0,CONN14,CH14,P1V8_DVDDIO_1,1.8
SWB0,CONN14,CH15,P0V75_DVDD_7,0.75
SWB0,CONN15,CH0,P1V5_AVDD_55,1.5
SWB0,CONN15,CH1,P0V9_AVDD_55,0.9
SWB0,CONN15,CH2,P0V75_DVDD_63,0.75
SWB0,CONN15,CH3,P0V75_AVDD_63,0.75
SWB0,CONN15,CH4,P1V5_AVDD_63,1.5
SWB0,CONN15,CH5,P0V9_AVDD_63,0.9
SWB0,CONN15,CH6,PVDD3_3V_SYNTH_LDO_B3,3.3
SWB0,CONN15,CH7,PVDD1V5_3,1.5
SWB0,CONN15,CH8,P0V85_STBY,0.85
SWB0,CONN15,CH9,P1V8_STBY,1.8
SWB0,CONN15,CH10,NC,-
SWB0,CONN15,CH11,P3V3_STBY,3.3
SWB0,CONN15,CH12,PVDD0V9_2,0.9
SWB0,CONN15,CH13,P0V75_DVDD_55,0.75
SWB0,CONN15,CH14,P0V75_AVDD_55,0.75
SWB0,CONN15,CH15,P1V8_DVDDIO_4,1.8
SWB0,CONN16,CH0,PVDD1V5_0,1.5
SWB0,CONN16,CH1,P1V8_4,1.8
SWB0,CONN16,CH2,PVDD3_3V_SYNTH_LDO_B5,3.3
SWB0,CONN16,CH3,PVDD3_3V_SYNTH_LDO_A5,3.3
SWB0,CONN16,CH4,PVDD3_3V_SYNTH_LDO_B4,3.3
SWB0,CONN16,CH5,PVDD3_3V_SYNTH_LDO_A4,3.3
SWB0,CONN16,CH6,P3V3,3.3
SWB0,CONN16,CH7,PVDD0_8V_T3,0.8
SWB0,CONN16,CH8,P0V75_DVDD_1,0.75
SWB0,CONN16,CH9,P0V75_AVDD_1,0.75
SWB0,CONN16,CH10,P1V5_AVDD_1,1.5
SWB0,CONN16,CH11,P0V9_AVDD_1,0.9
SWB0,CONN16,CH12,PVDD0V9_0,0.9
SWB0,CONN16,CH13,PVDD0V9_1,0.9
SWB0,CONN16,CH14,PVDD0V9_3,0.9
SWB0,CONN16,CH15,PVDD1V2_T3,1.2
SWB1,CONN13,CH0,P0V75_DVDD_47,0.75
SWB1,CONN13,CH1,P0V75_AVDD_47,0.75
SWB1,CONN13,CH2,P1V5_AVDD_47,1.5
SWB1,CONN13,CH3,P0V9_AVDD_47,0.9
SWB1,CONN13,CH4,P1V8_1,1.8
SWB1,CONN13,CH5,PVDD3_3V_SYNTH_LDO_A3,3.3
SWB1,CONN13,CH6,PVDD3_3V_SYNTH_LDO_B1,3.3
SWB1,CONN13,CH7,P1V8_3,1.8
SWB1,CONN13,CH8,P1V8_DVDDIO_3,1.8
SWB1,CONN13,CH9,P0V75_DVDD_39,0.75
SWB1,CONN13,CH10,P0V75_AVDD_39,0.75
SWB1,CONN13,CH11,P1V5_AVDD_39,1.5
SWB1,CONN13,CH12,P0V9_AVDD_39,0.9
SWB1,CONN13,CH13,PVDD1V5_2,1.5
SWB1,CONN13,CH14,PVDD1V8_DUT,1.8
SWB1,CONN13,CH15,PVDD_MDIO,1.2
SWB1,CONN14,CH0,P0V75_AVDD_7,0.75
SWB1,CONN14,CH1,P1V5_AVDD_7,1.5
SWB1,CONN14,CH2,P0V9_AVDD_7,0.9
SWB1,CONN14,CH3,P1V8_5,1.8
SWB1,CONN14,CH4,P1V8_2,1.8
SWB1,CONN14,CH5,PVDD3_3V_SYNTH_LDO_B2,3.3
SWB1,CONN14,CH6,PVDD3_3V_SYNTH_LDO_A2,3.3
SWB1,CONN14,CH7,PVDD3_3V_SYNTH_LDO_A1,3.3
SWB1,CONN14,CH8,PVDD1V5_TSC,1.5
SWB1,CONN14,CH9,PVDD1V5_ANLG,1.5
SWB1,CONN14,CH10,NC,-
SWB1,CONN14,CH11,NC,-
SWB1,CONN14,CH12,PVDD1V5_1,1.5
SWB1,CONN14,CH13,NC,-
SWB1,CONN14,CH14,P1V8_DVDDIO_1,1.8
SWB1,CONN14,CH15,P0V75_DVDD_7,0.75
SWB1,CONN15,CH0,P1V5_AVDD_55,1.5
SWB1,CONN15,CH1,P0V9_AVDD_55,0.9
SWB1,CONN15,CH2,P0V75_DVDD_63,0.75
SWB1,CONN15,CH3,P0V75_AVDD_63,0.75
SWB1,CONN15,CH4,P1V5_AVDD_63,1.5
SWB1,CONN15,CH5,P0V9_AVDD_63,0.9
SWB1,CONN15,CH6,PVDD3_3V_SYNTH_LDO_B3,3.3
SWB1,CONN15,CH7,PVDD1V5_3,1.5
SWB1,CONN15,CH8,P0V85_STBY,0.85
SWB1,CONN15,CH9,P1V8_STBY,1.8
SWB1,CONN15,CH10,NC,-
SWB1,CONN15,CH11,P3V3_STBY,3.3
SWB1,CONN15,CH12,PVDD0V9_2,0.9
SWB1,CONN15,CH13,P0V75_DVDD_55,0.75
SWB1,CONN15,CH14,P0V75_AVDD_55,0.75
SWB1,CONN15,CH15,P1V8_DVDDIO_4,1.8
SWB1,CONN16,CH0,PVDD1V5_0,1.5
SWB1,CONN16,CH1,P1V8_4,1.8
SWB1,CONN16,CH2,PVDD3_3V_SYNTH_LDO_B5,3.3
SWB1,CONN16,CH3,PVDD3_3V_SYNTH_LDO_A5,3.3
SWB1,CONN16,CH4,PVDD3_3V_SYNTH_LDO_B4,3.3
SWB1,CONN16,CH5,PVDD3_3V_SYNTH_LDO_A4,3.3
SWB1,CONN16,CH6,P3V3,3.3
SWB1,CONN16,CH7,PVDD0_8V_T3,0.8
SWB1,CONN16,CH8,P0V75_DVDD_1,0.75
SWB1,CONN16,CH9,P0V75_AVDD_1,0.75
SWB1,CONN16,CH10,P1V5_AVDD_1,1.5
SWB1,CONN16,CH11,P0V9_AVDD_1,0.9
SWB1,CONN16,CH12,PVDD0V9_0,0.9
SWB1,CONN16,CH13,PVDD0V9_1,0.9
SWB1,CONN16,CH14,PVDD0V9_3,0.9
SWB1,CONN16,CH15,PVDD1V2_T3,1.2
1 Board CONN Ch NetName Vnom
2 CB CONN13 CH0 V5P0_ALW 5.0
3 CB CONN13 CH1 V3P3_ALW 3.3
4 CB CONN13 CH2 PWR_VDD_MISC_ALW 0.75
5 CB CONN13 CH3 V1P8_ALW 1.8
6 CB CONN13 CH4 PWR_APU_VDDIO_SUS 1.1
7 CB CONN13 CH5 PWR_VDD_MISC_RUN 0.75
8 CB CONN13 CH6 PWR_APU_VDD_MEM_RUN 0.78
9 CB CONN13 CH7 V0P8_PHY2_DVDD 0.8
10 CB CONN13 CH8 V0P8_AVDD 0.8
11 CB CONN13 CH9 V0P8_PHY 0.8
12 CB CONN13 CH10 V1P8_FPGA 1.8
13 CB CONN13 CH11 V2P5_FPGA 2.5
14 CB CONN13 CH12 V1P1_FPGA 1.1
15 CB CONN13 CH13 P5V_STBY 5.0
16 CB CONN13 CH14 P3V3_STBY 3.3
17 CB CONN13 CH15 V0P8_PHY2_AVDD 0.8
18 SWB0 CONN13 CH0 P0V75_DVDD_47 0.75
19 SWB0 CONN13 CH1 P0V75_AVDD_47 0.75
20 SWB0 CONN13 CH2 P1V5_AVDD_47 1.5
21 SWB0 CONN13 CH3 P0V9_AVDD_47 0.9
22 SWB0 CONN13 CH4 P1V8_1 1.8
23 SWB0 CONN13 CH5 PVDD3_3V_SYNTH_LDO_A3 3.3
24 SWB0 CONN13 CH6 PVDD3_3V_SYNTH_LDO_B1 3.3
25 SWB0 CONN13 CH7 P1V8_3 1.8
26 SWB0 CONN13 CH8 P1V8_DVDDIO_3 1.8
27 SWB0 CONN13 CH9 P0V75_DVDD_39 0.75
28 SWB0 CONN13 CH10 P0V75_AVDD_39 0.75
29 SWB0 CONN13 CH11 P1V5_AVDD_39 1.5
30 SWB0 CONN13 CH12 P0V9_AVDD_39 0.9
31 SWB0 CONN13 CH13 PVDD1V5_2 1.5
32 SWB0 CONN13 CH14 PVDD1V8_DUT 1.8
33 SWB0 CONN13 CH15 PVDD_MDIO 1.2
34 SWB0 CONN14 CH0 P0V75_AVDD_7 0.75
35 SWB0 CONN14 CH1 P1V5_AVDD_7 1.5
36 SWB0 CONN14 CH2 P0V9_AVDD_7 0.9
37 SWB0 CONN14 CH3 P1V8_5 1.8
38 SWB0 CONN14 CH4 P1V8_2 1.8
39 SWB0 CONN14 CH5 PVDD3_3V_SYNTH_LDO_B2 3.3
40 SWB0 CONN14 CH6 PVDD3_3V_SYNTH_LDO_A2 3.3
41 SWB0 CONN14 CH7 PVDD3_3V_SYNTH_LDO_A1 3.3
42 SWB0 CONN14 CH8 PVDD1V5_TSC 1.5
43 SWB0 CONN14 CH9 PVDD1V5_ANLG 1.5
44 SWB0 CONN14 CH10 NC -
45 SWB0 CONN14 CH11 NC -
46 SWB0 CONN14 CH12 PVDD1V5_1 1.5
47 SWB0 CONN14 CH13 NC -
48 SWB0 CONN14 CH14 P1V8_DVDDIO_1 1.8
49 SWB0 CONN14 CH15 P0V75_DVDD_7 0.75
50 SWB0 CONN15 CH0 P1V5_AVDD_55 1.5
51 SWB0 CONN15 CH1 P0V9_AVDD_55 0.9
52 SWB0 CONN15 CH2 P0V75_DVDD_63 0.75
53 SWB0 CONN15 CH3 P0V75_AVDD_63 0.75
54 SWB0 CONN15 CH4 P1V5_AVDD_63 1.5
55 SWB0 CONN15 CH5 P0V9_AVDD_63 0.9
56 SWB0 CONN15 CH6 PVDD3_3V_SYNTH_LDO_B3 3.3
57 SWB0 CONN15 CH7 PVDD1V5_3 1.5
58 SWB0 CONN15 CH8 P0V85_STBY 0.85
59 SWB0 CONN15 CH9 P1V8_STBY 1.8
60 SWB0 CONN15 CH10 NC -
61 SWB0 CONN15 CH11 P3V3_STBY 3.3
62 SWB0 CONN15 CH12 PVDD0V9_2 0.9
63 SWB0 CONN15 CH13 P0V75_DVDD_55 0.75
64 SWB0 CONN15 CH14 P0V75_AVDD_55 0.75
65 SWB0 CONN15 CH15 P1V8_DVDDIO_4 1.8
66 SWB0 CONN16 CH0 PVDD1V5_0 1.5
67 SWB0 CONN16 CH1 P1V8_4 1.8
68 SWB0 CONN16 CH2 PVDD3_3V_SYNTH_LDO_B5 3.3
69 SWB0 CONN16 CH3 PVDD3_3V_SYNTH_LDO_A5 3.3
70 SWB0 CONN16 CH4 PVDD3_3V_SYNTH_LDO_B4 3.3
71 SWB0 CONN16 CH5 PVDD3_3V_SYNTH_LDO_A4 3.3
72 SWB0 CONN16 CH6 P3V3 3.3
73 SWB0 CONN16 CH7 PVDD0_8V_T3 0.8
74 SWB0 CONN16 CH8 P0V75_DVDD_1 0.75
75 SWB0 CONN16 CH9 P0V75_AVDD_1 0.75
76 SWB0 CONN16 CH10 P1V5_AVDD_1 1.5
77 SWB0 CONN16 CH11 P0V9_AVDD_1 0.9
78 SWB0 CONN16 CH12 PVDD0V9_0 0.9
79 SWB0 CONN16 CH13 PVDD0V9_1 0.9
80 SWB0 CONN16 CH14 PVDD0V9_3 0.9
81 SWB0 CONN16 CH15 PVDD1V2_T3 1.2
82 SWB1 CONN13 CH0 P0V75_DVDD_47 0.75
83 SWB1 CONN13 CH1 P0V75_AVDD_47 0.75
84 SWB1 CONN13 CH2 P1V5_AVDD_47 1.5
85 SWB1 CONN13 CH3 P0V9_AVDD_47 0.9
86 SWB1 CONN13 CH4 P1V8_1 1.8
87 SWB1 CONN13 CH5 PVDD3_3V_SYNTH_LDO_A3 3.3
88 SWB1 CONN13 CH6 PVDD3_3V_SYNTH_LDO_B1 3.3
89 SWB1 CONN13 CH7 P1V8_3 1.8
90 SWB1 CONN13 CH8 P1V8_DVDDIO_3 1.8
91 SWB1 CONN13 CH9 P0V75_DVDD_39 0.75
92 SWB1 CONN13 CH10 P0V75_AVDD_39 0.75
93 SWB1 CONN13 CH11 P1V5_AVDD_39 1.5
94 SWB1 CONN13 CH12 P0V9_AVDD_39 0.9
95 SWB1 CONN13 CH13 PVDD1V5_2 1.5
96 SWB1 CONN13 CH14 PVDD1V8_DUT 1.8
97 SWB1 CONN13 CH15 PVDD_MDIO 1.2
98 SWB1 CONN14 CH0 P0V75_AVDD_7 0.75
99 SWB1 CONN14 CH1 P1V5_AVDD_7 1.5
100 SWB1 CONN14 CH2 P0V9_AVDD_7 0.9
101 SWB1 CONN14 CH3 P1V8_5 1.8
102 SWB1 CONN14 CH4 P1V8_2 1.8
103 SWB1 CONN14 CH5 PVDD3_3V_SYNTH_LDO_B2 3.3
104 SWB1 CONN14 CH6 PVDD3_3V_SYNTH_LDO_A2 3.3
105 SWB1 CONN14 CH7 PVDD3_3V_SYNTH_LDO_A1 3.3
106 SWB1 CONN14 CH8 PVDD1V5_TSC 1.5
107 SWB1 CONN14 CH9 PVDD1V5_ANLG 1.5
108 SWB1 CONN14 CH10 NC -
109 SWB1 CONN14 CH11 NC -
110 SWB1 CONN14 CH12 PVDD1V5_1 1.5
111 SWB1 CONN14 CH13 NC -
112 SWB1 CONN14 CH14 P1V8_DVDDIO_1 1.8
113 SWB1 CONN14 CH15 P0V75_DVDD_7 0.75
114 SWB1 CONN15 CH0 P1V5_AVDD_55 1.5
115 SWB1 CONN15 CH1 P0V9_AVDD_55 0.9
116 SWB1 CONN15 CH2 P0V75_DVDD_63 0.75
117 SWB1 CONN15 CH3 P0V75_AVDD_63 0.75
118 SWB1 CONN15 CH4 P1V5_AVDD_63 1.5
119 SWB1 CONN15 CH5 P0V9_AVDD_63 0.9
120 SWB1 CONN15 CH6 PVDD3_3V_SYNTH_LDO_B3 3.3
121 SWB1 CONN15 CH7 PVDD1V5_3 1.5
122 SWB1 CONN15 CH8 P0V85_STBY 0.85
123 SWB1 CONN15 CH9 P1V8_STBY 1.8
124 SWB1 CONN15 CH10 NC -
125 SWB1 CONN15 CH11 P3V3_STBY 3.3
126 SWB1 CONN15 CH12 PVDD0V9_2 0.9
127 SWB1 CONN15 CH13 P0V75_DVDD_55 0.75
128 SWB1 CONN15 CH14 P0V75_AVDD_55 0.75
129 SWB1 CONN15 CH15 P1V8_DVDDIO_4 1.8
130 SWB1 CONN16 CH0 PVDD1V5_0 1.5
131 SWB1 CONN16 CH1 P1V8_4 1.8
132 SWB1 CONN16 CH2 PVDD3_3V_SYNTH_LDO_B5 3.3
133 SWB1 CONN16 CH3 PVDD3_3V_SYNTH_LDO_A5 3.3
134 SWB1 CONN16 CH4 PVDD3_3V_SYNTH_LDO_B4 3.3
135 SWB1 CONN16 CH5 PVDD3_3V_SYNTH_LDO_A4 3.3
136 SWB1 CONN16 CH6 P3V3 3.3
137 SWB1 CONN16 CH7 PVDD0_8V_T3 0.8
138 SWB1 CONN16 CH8 P0V75_DVDD_1 0.75
139 SWB1 CONN16 CH9 P0V75_AVDD_1 0.75
140 SWB1 CONN16 CH10 P1V5_AVDD_1 1.5
141 SWB1 CONN16 CH11 P0V9_AVDD_1 0.9
142 SWB1 CONN16 CH12 PVDD0V9_0 0.9
143 SWB1 CONN16 CH13 PVDD0V9_1 0.9
144 SWB1 CONN16 CH14 PVDD0V9_3 0.9
145 SWB1 CONN16 CH15 PVDD1V2_T3 1.2
+56
View File
@@ -0,0 +1,56 @@
# v1.0.10 — The soak loop stops spinning flat out, and shows you what is still running
## ✨ New features
**Every soak round now reports the background jobs**
* Script B's monitoring loop prints `bgctl list` (platform jobs) and `jobs` (the login shell's own background jobs) on every pass, right after the platform data. A stress job that dies three hours into an overnight soak now shows up as a shorter list on the next round, instead of only being noticed at the end — or not at all.
* The two lists answer different questions and are both worth having: `bgctl` knows about the platform's job runner, `jobs` knows about anything the macro backgrounded in that shell.
**A single table for all 144 margin channels**
* `docs/LTC2980_channel_map.csv` flattens the nine `settings/*.conf` files into one sheet — `Board,CONN,Ch,NetName,Vnom` — so "which board and channel is `PVDD1V5_TSC` on?" is one search instead of opening nine files.
* `tools/gen_channel_map.sh` regenerates it. **The CSV is output, not source**: edit the `.conf` files and re-run the script, never the other way round.
* Unused channels stay in the table as `NC` / `-` rather than being dropped, because the channel number is also the PMBus page (`ch / 8` picks the chip, `ch % 8` picks the page) — filtering the gaps out would shift everything after them.
**The bench command lists that these scripts came from**
* `tools/100G_PRBS.txt`, `tools/TR518.txt` and `tools/fan_ctrl.txt` join the existing `traffic_loopback_*.txt` set: the raw `bcmcmd` / sysfs sequences, kept in the form you can paste into a console when a script is misbehaving and you want to drive the hardware directly.
## 🐛 Bug fixes
**The soak loop ran as fast as the DUT could answer**
* `while 1` had no pause at all, so each round started the moment the previous one finished. Over a long soak that means constant console traffic and a log full of near-identical samples taken seconds apart. The loop now ends each round with `pause 60`.
* Note the gap this leaves: the Status sheet asks for a 10-minute sampling interval, and 60 seconds is still hard-coded rather than configurable. This makes the loop sane, it does not finish the job.
**A comment promised a cadence the code never had**
* The loop was labelled `Get data every 10mins` while running with no delay whatsoever. It now says what it does.
## 📦 Downloads
| File | Contents |
|---|---|
| `Script_ABC_Blanton_V1.0.10.zip` | The full Tera Term working directory, including `Blanton_Script/` to copy onto the DUT |
**After copying to the DUT:**
```bash
chmod +x ~/Blanton_Script/*.sh # see below
grep -rlU $'\r' ~/Blanton_Script # expect no output
```
`mgmt_ping_monitor.sh` and `port_prbs_monitor.sh` launch their workers through `bash`, so those two run without the execute bit. `bmc_monitor.sh` and `usb_target.sh` still need it, and neither git (mode 100644) nor a Windows/USB copy carries it. Without the `chmod`, `start` fails with `Permission denied` and the later `cat` finds nothing: **the section ends up empty and nothing reports an error.**
## ⚠️ Before you run
* **Declare the switch population.** `SWB_UNIT0` / `SWB_UNIT1` in `config.ttl` say which switch units this DUT has; traffic and the readiness gate both follow them.
* **Set `FAN_SPEED`.** It is applied at the start of every run.
* **Management connectivity is in use, not preserved.** Both management NICs are up and pinging — drive the run from the serial console.
* **PRBS is still not wired into the macros.** Run `port_prbs_monitor.sh` by hand if you want it; the calls in A/B/C remain commented out, unchanged from v1.0.9.
* **TR518 and the loopback traffic test cannot both run.** `tools/TR518.txt` sets `port all lb=mac` and `l2 learn off` across the whole unit, which is the same hardware `blanton_traffic_linespeed` is using. Pick one.
* **Two settings persist to `config_db.json`** and survive a reboot: LLDP is left disabled and the 100G uplinks are left configured. Restore them before the DUT moves on.
## 🔗 Links
* [ARCHITECTURE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/ARCHITECTURE.md) — test flow, data models, constraints
* [CLAUDE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/CLAUDE.md) — commands and the bench gotchas
* [LTC2980_channel_map.csv](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/docs/LTC2980_channel_map.csv) — all 144 margin channels
**Full changelog:** [V1.0.9...V1.0.10](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/compare/V1.0.9...V1.0.10)
+59
View File
@@ -0,0 +1,59 @@
# v1.0.11 — Saving the margin settings to NVM is now one command, and the command is correct
## ✨ New features
**`margin_save blanton` — all 18 LTC2977 in one go**
* On this platform the CB FPGA's I2C channels are also exposed as ordinary Linux i2c buses, so every margin controller on the box — CB and both switch boards — can be reached with plain `i2cset`. `margin_save blanton` walks all 18 of them from one flat list: no `margin_init`, no FPGA channel setup, no `pcimem`.
* The channel-to-bus mapping it relies on (F3 Ch6 → bus 13, Ch7 → 14, Ch8 → 15, Ch9 → 16; CB is bus 4 as before) is now written down in `ARCHITECTURE.md` and the margin README, not just known at the bench.
* **A partly-populated DUT is a normal build.** Comment out the lines you do not have in `MARGIN_BLANTON_STORE` at the top of `margin.sh` rather than letting the tool talk to boards that are not there.
* Missing bus and unresponsive chip are reported as two different failures. `/dev/i2c-16 does not exist` means the FPGA's i2c adapters did not enumerate; `i2cset failed (NACK/busy)` means the board is absent or the address is wrong. Chasing the second when you have the first wastes an afternoon.
**`margin_save all` — every board in `settings/`**
* Runs `margin_init` + `margin_save` across all nine boards, then puts back whichever board you had loaded, so a following `margin_status` still reports the board you were looking at.
* It does not stop at the first failure: a half-saved set is harder to reason about than a fully-attempted one. The summary names the boards that failed.
Both print a per-chip line and a `N saved, M failed` tally, and both return a non-zero exit code when anything failed.
**The fans now default to full speed**
* `FAN_SPEED` in `config.ttl` is 100 instead of 30. A soak is a thermal test of everything *except* the cooling, so the cooling should not be a variable in it — and a DUT that throttles halfway through a run produces results that look like a different fault. Lower it deliberately if quiet operation is what you are measuring.
## 🐛 Bug fixes
**A failed NVM write still reported success**
* `margin_save` discarded the result of every write, so a NACK printed `Change Saved` exactly like a successful store. For the one operation in this tool that is permanent, that is the worst place to stay quiet. Each store is now checked and named if it fails.
**`STORE_USER_ALL` was sent with a data byte it should not have**
* The command went out as `0x15` followed by a dummy `0x00` — an SMBus write-byte, where `STORE_USER_ALL` is a send-byte. The LTC2977 tolerated it, but tolerance is not correctness. It is now `i2cset ... 0x15 c` on the native path and a no-data `cb_pmbus_write` on the FPGA path.
## 📦 Downloads
| File | Contents |
|---|---|
| `Script_ABC_Blanton_V1.0.11.zip` | The full Tera Term working directory, including `Blanton_Script/` to copy onto the DUT |
**After copying to the DUT:**
```bash
chmod +x ~/Blanton_Script/*.sh # see below
grep -rlU $'\r' ~/Blanton_Script # expect no output
```
`mgmt_ping_monitor.sh` and `port_prbs_monitor.sh` launch their workers through `bash`, so those two run without the execute bit. `bmc_monitor.sh` and `usb_target.sh` still need it, and neither git (mode 100644) nor a Windows/USB copy carries it. Without the `chmod`, `start` fails with `Permission denied` and the later `cat` finds nothing: **the section ends up empty and nothing reports an error.**
## ⚠️ Before you run
* **`margin_save` in any form writes NVM permanently.** A power cycle does not undo it. `margin_save all` and `margin_save blanton` do it to all 18 controllers at once. Run `margin_status_all.sh` first and never call any of them during a test run — the macros do not, and should not be made to.
* **The 18 raw commands were verified on the DUT; the wrapper was not.** `margin_save blanton` issues exactly the sequence that was run by hand at the bench, and the packaging was tested against stubbed hardware, but the function itself has not been run on a DUT yet. Watch the per-chip lines the first time.
* **Declare the switch population.** `SWB_UNIT0` / `SWB_UNIT1` in `config.ttl` say which switch units this DUT has; traffic and the readiness gate both follow them.
* **`FAN_SPEED` is now 100.** It is applied at the start of every run, so this build runs the fans flat out unless you change `config.ttl`. Expect the noise.
* **Management connectivity is in use, not preserved.** Both management NICs are up and pinging — drive the run from the serial console.
* **PRBS is still not wired into the macros.** Run `port_prbs_monitor.sh` by hand if you want it; the calls in A/B/C remain commented out.
* **Two settings persist to `config_db.json`** and survive a reboot: LLDP is left disabled and the 100G uplinks are left configured. Restore them before the DUT moves on.
## 🔗 Links
* [ARCHITECTURE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/ARCHITECTURE.md) — test flow, data models, constraints
* [CLAUDE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/CLAUDE.md) — commands and the bench gotchas
* [LTC2980_Margin_Script/README.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/src/Script_ABC_Blanton/Blanton_Script/LTC2980_Margin_Script/README.md) — the three `margin_save` variants and what each one sends
**Full changelog:** [V1.0.10...V1.0.11](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/compare/V1.0.10...V1.0.11)
+65
View File
@@ -0,0 +1,65 @@
# v1.0.7 — Management-link ping test, and the fans start where you set them
## ✨ New features
**10G / 1G management port ping test**
* New `mgmt_ping_monitor.sh` exercises both management NICs in one loop: it brings each up with iproute2, waits for the link to settle, then pings that NIC's own target — `eth0``192.168.1.30`, `eth1``192.168.1.31`.
* Runs detached like the other monitors: `start` / `stop` / `status` / `fg` / `summary`. `start` clears the previous log, so one run means one log.
* Every leg writes a single greppable line — `RESULT 10G eth0 -> 192.168.1.30 tx=10 rx=10 loss=0% rtt_avg=1.743ms nic_tx=10 nic_rx=10 PASS` — so a whole soak can be read with `summary` instead of scrolling.
* `nic_tx` / `nic_rx` are that interface's own counter delta across the burst. If a ping passes but they stay near zero, the traffic left on the *other* NIC and the number does not belong to this one.
* Script A runs one round as a baseline; Script B leaves it running through the soak.
**Fan speed is set at the start of every run**
* `FAN_SPEED` in `config.ttl` (default 30) is applied by Script A, so a run no longer inherits whatever the previous test left the fans doing.
* Both MAX31790 controllers are rebound first, which clears a controller left in an odd state.
**Stress load moved onto the platform's own job runner**
* CPU, DDR, SSD and USB stress now go through `bgctl`, replacing the `~/hammer/tools/stress_*.py` scripts. `bgctl list` shows what is running; Script C stops everything with `bgctl stop --all`.
* BMC DDR stress runs through `bmc-manager`, and a new BMC USB test discovers the `cdc_ncm` interface and pings across it for four hours under `systemd-run`.
**More of the DUT on record**
* Script A now captures the boot image, `show version`, firmware status, system EEPROM, SSD health, TPM version and NVMe SMART data before the test starts.
* PCIe AER now sits next to `ras-mc-ctl --summary`, and Script C closes with NVMe health so a disk that degraded during the soak is visible.
* `dmesg` capture uses human-readable timestamps and one combined error pattern, then clears the ring buffer — so Script C's dmesg shows only what the soak produced, not everything since boot.
* Each run opens with `hw-test-session start` and closes with `finish`; Script C copies the job logs to `/mnt/usb/jobs-<date>_<time>` so they leave the DUT with the run they belong to.
## 🐛 Bug fixes
**Script A no longer hangs waiting for ports that are already up**
* The readiness gate required *more than* 216 ports up, but a fully cabled unit reports exactly 216 — 108 loopback pairs times two. Script A sat in the poll loop forever on the very state it was waiting for. It now proceeds at 216.
* The comments in `wait_init.ttl` still described the old "more than" rule; they now match the code, so the off-by-one cannot be reintroduced by reading the wrong line.
**The macro no longer runs ahead of the DUT during the soak**
* Script B set a 15-second wait cap for the `cdc_ncm` probe and never restored it. In Tera Term that cap is global, so it stayed in force for everything after — including the traffic setup, which walks 108 VLANs per switch unit and takes far longer. A timed-out wait returns without the prompt, leaving every later command issued a step early.
**Fan controller is no longer bound twice**
* One of the two MAX31790 controllers received a second `bind` immediately after the first, which could only fail because the driver was already attached.
## 📦 Downloads
| File | Contents |
|---|---|
| `Script_ABC_Blanton_V1.0.7.zip` | The full Tera Term working directory, including `Blanton_Script/` to copy onto the DUT |
**After copying to the DUT:**
```bash
chmod +x ~/Blanton_Script/*.sh # required -- see below
grep -rlU $'\r' ~/Blanton_Script # expect no output
```
The `chmod +x` is not optional. `mgmt_ping_monitor.sh` is executed rather than sourced and re-execs its own path to daemonise, and neither git (mode 100644) nor a Windows/USB copy carries the execute bit. Without it `start` fails with `Permission denied` and the later `cat` finds nothing — **the section ends up empty and nothing reports an error**.
## ⚠️ Before you run
* **Declare the switch population.** `SWB_UNIT0` / `SWB_UNIT1` in `config.ttl` say which switch units this DUT has. Traffic and the readiness gate both follow them: a single unit gets `-u 0` or `-u 1`, and with neither set the traffic stage is skipped instead of failing against hardware that is not there.
* **Both management NICs stay up.** If they share a subnet, keep `ARP_STRICT=1` — otherwise the target's ARP can be answered by either NIC and a reply may arrive on the one that did not send.
* **Management connectivity is exercised, not preserved.** Drive the run from the serial console.
* **Two settings persist to `config_db.json`** and survive a reboot: LLDP is left disabled, and the 100G uplinks are left configured. Restore them before the DUT moves on.
## 🔗 Links
* [ARCHITECTURE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/ARCHITECTURE.md) — test flow, data models, constraints
* [CLAUDE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/CLAUDE.md) — commands and the bench gotchas
**Full changelog:** [V1.0.5...V1.0.7](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/compare/V1.0.5...V1.0.7)
+61
View File
@@ -0,0 +1,61 @@
# v1.0.8 — Both management links tested at once, and ARP stops looking like packet loss
## ✨ New features
**Both management NICs are pinged simultaneously and continuously**
* `mgmt_ping_monitor.sh` no longer alternates fixed bursts. `start` configures both NICs, then pings from each at the same time and keeps accumulating until stopped — so a soak-long run is one continuous measurement rather than a series of snapshots.
* `stop` renders a report into the log: the last 20 entries per NIC, each one's ping statistics and PASS/FAIL, then `ip -s link show` for both interfaces.
* Every line is timestamped, and a request that got no reply prints a marker instead of merely being absent — a drop is visible in the tail, not inferred from a gap in the sequence numbers.
* `status` shows both pids and how many replies each NIC has received so far, which is the quick way to see one side is dead without waiting for the report.
**The USB stress target is detected, not assumed**
* New `usb_target.sh` reports an inserted USB device's node, mount point or by-id path. Script B asks it for the node and builds the stress command from the answer, so a stick that enumerates as `sdb` no longer sends a write test at whatever `/dev/sda1` happens to be.
* It excludes whatever disk backs `/` and `/host` — this platform can boot from a USB DOM, which also reports as USB — refuses to guess when several USB disks are present, and when it mounts, verifies the result is actually writable rather than trusting that `mount` succeeded.
**BMC I2C integrity is exercised through the soak**
* `bmc_monitor.sh` is back in the run and now writes two complementary patterns to a BMC scratch register and reads each back. One pattern alone cannot catch a bit stuck the same way it was written; repeating the pair through the soak turns an intermittent I2C fault into something the log records rather than something the tester has to witness.
**Test artefacts leave the DUT with the run**
* Script C copies `bmc_poll.log` and `mgmt_ping.log` into the job directory before it is archived to USB, so the monitors' output travels with the `bgctl` job logs.
* DDR stress now runs continuously instead of stopping after 100 passes, matching the other soak loads.
## 🐛 Bug fixes
**Three FAILs that were not link faults**
* A bench run reported `FAIL=3` with zero NIC errors, zero drops and sub-millisecond replies. Every failure was missing exactly `icmp_seq=1` and nothing else: once the neighbour entry for the target expires, the first echo request is spent resolving ARP and `ping` counts it as loss. A discarded warm-up ping per NIC now absorbs that, which is what lets the loss threshold stay at zero and still mean something. Raising the threshold instead would have hidden genuine single-packet loss.
**Job logs were being cleared before they were collected**
* Script C ran `bgctl reset --yes` while killing processes — before the job directory is copied to USB. Moved to after the archive, so a run's own logs are collected before anything clears them.
**The fan controller was bound twice**
* One of the two MAX31790 controllers received a second `bind` immediately after the first, which could only fail because the driver was already attached.
## 📦 Downloads
| File | Contents |
|---|---|
| `Script_ABC_Blanton_V1.0.8.zip` | The full Tera Term working directory, including `Blanton_Script/` to copy onto the DUT |
**After copying to the DUT:**
```bash
chmod +x ~/Blanton_Script/*.sh # see below
grep -rlU $'\r' ~/Blanton_Script # expect no output
```
`mgmt_ping_monitor.sh` no longer re-execs itself, so `bash mgmt_ping_monitor.sh start` works without the execute bit. `bmc_monitor.sh` and `usb_target.sh` still need it — and neither git (mode 100644) nor a Windows/USB copy carries it. Without the `chmod`, `start` fails with `Permission denied` and the later `cat` finds nothing: **the section ends up empty and nothing reports an error.**
## ⚠️ Before you run
* **Declare the switch population.** `SWB_UNIT0` / `SWB_UNIT1` in `config.ttl` say which switch units this DUT has. Traffic and the readiness gate both follow them; with neither set the traffic stage is skipped rather than failing against hardware that is not there.
* **Set the fan speed you want.** `FAN_SPEED` in `config.ttl` is applied at the start of every run.
* **Management connectivity is in use, not preserved.** Both NICs are up and pinging — drive the run from the serial console.
* **If the two management NICs share a subnet, keep `ARP_STRICT=1`.** Otherwise the target's ARP can be answered by either NIC and a reply may arrive on the one that did not send. Check `nic_tx` / `nic_rx` on the RESULT line: they are that interface's own counter delta.
* **Two settings persist to `config_db.json`** and survive a reboot: LLDP is left disabled and the 100G uplinks are left configured. Restore them before the DUT moves on.
## 🔗 Links
* [ARCHITECTURE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/ARCHITECTURE.md) — test flow, data models, constraints
* [CLAUDE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/CLAUDE.md) — commands and the bench gotchas
**Full changelog:** [V1.0.7...V1.0.8](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/compare/V1.0.7...V1.0.8)
+63
View File
@@ -0,0 +1,63 @@
# v1.0.9 — 100G PRBS testing, and the job logs actually reach the USB stick
## ✨ New features
**100G uplink PRBS test**
* New `port_prbs_monitor.sh` drives a PRBS test on both 100G uplinks: it clears low-power mode, arms the pattern, then polls `phy diag <phy> prbs get` in the background until stopped. A poll counts as PASS only when the output says `PRBS OK!`; `prbsstat Ber` is captured alongside every poll for the record.
* `stop` runs `prbsstat STOp` and `prbs clear` on each port before rendering the report, so the test does not leave PRBS armed behind it.
* `report` gives the tally the bench actually wants:
```
PORT POLLS PASS FAIL RESULT
Ethernet513 42 42 0 PASS
Ethernet514 42 40 2 FAIL
```
* **One port at a time is a first-class mode.** `start a` or `start b` runs a single uplink, which is how you separate a genuine port fault from the DUT not coping with two PRBS streams at once. A port that was not run reports `SKIP`, not `FAIL`.
* Every command and its full output — setup, each poll, teardown — goes to that port's raw log.
> ⚠️ The script ships in this build and can be run by hand, but **the calls in Script A, B and C are commented out**: PRBS is still under bring-up on this platform.
**USB target detection**
* `usb_target.sh` reports an inserted USB device's node, mount point or by-id path. It excludes whatever disk backs `/` and `/host` — this platform can boot from a USB DOM, which also reports as USB — refuses to guess when several USB disks are present, and when mounting, verifies the result is writable rather than trusting that `mount` succeeded.
* Script B builds the USB stress command from the detected node instead of a hard-coded `/dev/sda1`.
## 🐛 Bug fixes
**The job logs were being written to a path nobody collected**
* Stress output, the monitors' logs and the USB archive disagreed about where the job directory lives. Everything now uses `/host/hw-eval/jobs/`, so what Script B writes is what Script C copies out.
**Script C copied to a USB that was never mounted**
* The archive step assumed `/mnt/usb` was ready. Script C now finds the device and mounts it first, so a run's logs leave the DUT instead of being written into an empty mount point.
**`show interfaces status` right after starting PRBS reads as a failure**
* With PRBS armed the link is out of normal operation and reports DOWN. `port_prbs_monitor.sh` deliberately does not print the port status at `start` — it appears in the report instead, after `prbs clear`, labelled as the recovered state.
## 📦 Downloads
| File | Contents |
|---|---|
| `Script_ABC_Blanton_V1.0.9.zip` | The full Tera Term working directory, including `Blanton_Script/` to copy onto the DUT |
**After copying to the DUT:**
```bash
chmod +x ~/Blanton_Script/*.sh # see below
grep -rlU $'\r' ~/Blanton_Script # expect no output
```
`mgmt_ping_monitor.sh` and `port_prbs_monitor.sh` launch their workers through `bash`, so those two run without the execute bit. `bmc_monitor.sh` and `usb_target.sh` still need it, and neither git (mode 100644) nor a Windows/USB copy carries it. Without the `chmod`, `start` fails with `Permission denied` and the later `cat` finds nothing: **the section ends up empty and nothing reports an error.**
## ⚠️ Before you run
* **Declare the switch population.** `SWB_UNIT0` / `SWB_UNIT1` in `config.ttl` say which switch units this DUT has; traffic and the readiness gate both follow them.
* **Set `FAN_SPEED`.** It is applied at the start of every run.
* **Management connectivity is in use, not preserved.** Both management NICs are up and pinging — drive the run from the serial console.
* **PRBS is not wired into the macros yet.** Run `port_prbs_monitor.sh` by hand if you want it; uncommenting the calls in A/B/C is not supported in this build.
* **Two settings persist to `config_db.json`** and survive a reboot: LLDP is left disabled and the 100G uplinks are left configured. Restore them before the DUT moves on.
## 🔗 Links
* [ARCHITECTURE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/ARCHITECTURE.md) — test flow, data models, constraints
* [CLAUDE.md](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/src/branch/main/CLAUDE.md) — commands and the bench gotchas
**Full changelog:** [V1.0.8...V1.0.9](https://etgit.et-wen.com/etwen/Blanton_TTL_Script/compare/V1.0.8...V1.0.9)
+259 -26
View File
@@ -1,7 +1,7 @@
; =============================================================================
; Script A for Blanton
; Version : V1.0.3
; Date : 2026-08-21
; Version : V1.0.11
; Date : 2026-08-28
; Author : ETWen
; =============================================================================
; Version History:
@@ -16,6 +16,85 @@
; ScriptB/C Add Traffic Test
; V1.0.3 2026-08-21 utils/wait_init.ttl Created
; ScriptA Add Wait DUT Ready
; V1.0.4 2026-08-21 utils/wait_init.ttl Wait on bcmcmd port-up, count (both units > 216)
; utils/wait_init.ttl Add WT_UNIT0/WT_UNIT1, skip absent units
; ScriptA Add EEPROM info (hpe-eeprom-tlv --bus 4 --addr 0x50)
; ScriptA Remove empty Check History block
; ScriptA/B Disable lldp + config save before traffic (silent MAC)
; ScriptB Add sfputil lpmode off Ethernet513/514
; V1.0.5 2026-08-21 config.ttl Add SWB_UNIT0/SWB_UNIT1
; utils/wait_init.ttl Follow config.ttl SWB_UNIT0/SWB_UNIT1
; ScriptA/B/C Traffic follows SWB_UNIT: -u 0 / -u 1 / skip
; V1.0.6 2026-08-24 utils/show_dmesg.ttl Edited dmesg grep i2c & clear event
; ScriptA HW Test Session
; ScriptC HW Test Session Finish
; ScriptA bmc-first-enroll, bmc version/status
; ScriptB Stress test DDR, SSD, USB, BMC DDR, BMC USB Stress test
; ScriptC BMC USB Stress test result
; ScriptA DUT Info
; ScriptAC 10G/1G MGMT ping
; ScriptA Edited 100G Port Status
; ScriptA, utils/setup_pmon.ttl Add TPM
; ScriptAC NVME Error Info
; ScriptAC Add PCIe Summary
; V1.0.7 2026-08-25 config.ttl Add FAN_SPEED
; ScriptA Add Fan Ctrl (max31790 rebind + fan-speed-control.sh)
; ScriptA Clear /host/hw-eval/current/jobs before the run
; ScriptC Copy stress jobs log to USB, timestamped
; utils/wait_init.ttl WT_INTERVAL 10 -> 60, fix WT_MIN comments
; ScriptC Disable BMC USB journalctl dump
; V1.0.8 2026-08-25 Blanton_Script/bmc_monitor.sh Add BMC I2C write/read-back pattern test
; ScriptB Re-enable bmc_monitor.sh start
; ScriptC Re-enable bmc_monitor.sh stop + cat bmc_poll.log
; ScriptC Copy bmc_poll.log + mgmt_ping.log into jobs/ before the USB archive
; Blanton_Script/usb_target.sh Created - auto-detect USB device node / mount point
; ScriptB USB stress uses the detected node, not a hard-coded /dev/sda1
; ScriptB DDR memtester runs continuously (drop the 100-pass limit)
; ScriptC Move bgctl reset --yes after the USB archive
; Blanton_Script/mgmt_ping_monitor.sh V3.0.0 - both NICs ping simultaneously, accumulating
; + ARP warm-up so a stale neighbour entry is not counted as loss
; + stop renders last-N per NIC, stats, and both ip -s link show
; .gitignore Ignore monitor *.raw / *.pid / *.state
; ScriptA MGMT ping baseline 45s -> 30s, clear the log afterwards
; ScriptB Show mgmt_ping status after start
; ScriptC Show mgmt_ping status before stop, cat the report after
; V1.0.9 2026-08-26 Blanton_Script/port_prbs_monitor.sh Add 100G Port PRBS test
; start [both|a|b] - single port isolates a port fault from
; the DUT not coping with two PRBS streams at once
; !! TTL calls are commented out in A/B/C - still under bring-up
; ScriptA/B/C Job log path /host/hw-eval/current/jobs -> /host/hw-eval/jobs
; ScriptC Mount the USB (usb_target.sh) before copying the job logs
; ScriptC Reorder results: stress logs, then MGMT ping, then 100G status
; V1.0.10 2026-08-27 ScriptB Monitor loop lists the background jobs every round
; + bgctl list - platform jobs
; + jobs - shell jobs of the login shell
; + pause 60 - the loop used to run with no pause at all
; docs/LTC2980_channel_map.csv Channel / net / Vnom map built from settings/*.conf
; tools/gen_channel_map.sh Regenerates that CSV - the CSV is output, not source
; tools/100G_PRBS.txt Bench command references, kept for hand-run debug
; tools/TR518.txt TR518 = built-in packet test (swutil / bcmcmd tr 518)
; tools/fan_ctrl.txt MAX31790 rebind + fan-speed-control.sh
; V1.0.11 2026-08-28 LTC2980_Margin_Script/margin.sh v2.7.0 - margin_save fixes + two new subcommands
; STORE_USER_ALL (0x15) is now a PMBus send-byte, not
; 0x15 followed by a dummy 0x00 data byte. The LTC2977
; tolerated the old form, but it was the wrong protocol
; i2c: i2cset .. 0x15 c swb: cb_pmbus_write <ch> <addr> 0x15
; + margin_save all - every board in settings/ (9 boards,
; 18 chips), then re-init the one loaded beforehand
; + margin_save blanton - Blanton SONiC shortcut: all 18
; LTC2977 straight over the native i2c buses - no
; margin_init, no FPGA channel setup, no pcimem
; F3 Ch6->bus 13, Ch7->14, Ch8->15, Ch9->16
; board list = MARGIN_BLANTON_STORE at the top of margin.sh
; The store result is now checked - a NACK used to still
; print "Change Saved"
; MARGIN_STORE_SETTLE (0.5s) between stores; nothing in this
; tool polls the part's busy bit
; LTC2980_Margin_Script/README.md What each margin_save variant sends, and the
; FPGA-channel -> i2c-bus map
; LTC2980_Margin_Script/Ref/*.md 0x15 trace note corrected to send-byte
; LTC2980_Margin_Script/script/margin_save_all.sh Point at margin_save all; keep it for subsets
; config.ttl FAN_SPEED 30 -> 100 (full speed by default)
; =============================================================================
include "config.ttl"
@@ -50,17 +129,75 @@ gettime timestr
sprintf2 cmd 'sudo date -s "%s %s"' datestr timestr
sendln cmd
; ========== Show Build Name ==========
;Stress log clear
wait prompt_sonic_root
sendln "show boot"
sendln "rm /host/hw-eval/jobs/*"
; ========== HW Test Session ==========
wait prompt_sonic_root
sendln "hw-test-session start"
wait prompt_sonic_root
sendln "hw-test-session log"
wait prompt_sonic_root
sendln "hw-test-session status"
; ========== FAN SPEED ==========
wait prompt_sonic_root
sendln "echo 18-0020 > /sys/bus/i2c/drivers/max31790/unbind"
wait prompt_sonic_root
sendln "sleep .5"
wait prompt_sonic_root
sendln "echo 18-0020 > /sys/bus/i2c/drivers/max31790/bind"
wait prompt_sonic_root
sendln "echo 25-0020 > /sys/bus/i2c/drivers/max31790/unbind"
wait prompt_sonic_root
sendln "sleep .5"
wait prompt_sonic_root
sendln "echo 25-0020 > /sys/bus/i2c/drivers/max31790/bind"
wait prompt_sonic_root
sprintf2 fan_ctrl_cmd 'fan-speed-control.sh %d' FAN_SPEED
sendln fan_ctrl_cmd
wait "Set all configured fan channels to"
sendln "y"
; ========== Wait DUT Init ==========
wait prompt_sonic_root
include "utils/wait_init.ttl"
; ========== BMC version ==========
; ========== DUT Info ==========
wait prompt_sonic_root
sendln "bmc-manager run 'cat /etc/issue'"
sendln "show boot"
wait prompt_sonic_root
sendln "show version"
wait prompt_sonic_root
sendln "fwutil show status"
wait prompt_sonic_root
sendln "show system-memory"
wait prompt_sonic_root
sendln "show service"
wait prompt_sonic_root
sendln "show platform syseeprom"
wait prompt_sonic_root
sendln "show platform ssdhealth"
wait prompt_sonic_root
sendln "qfx5252-tpm-version"
wait prompt_sonic_root
sendln "/usr/sbin/nvme smart-log -H /dev/nvme0"
wait prompt_sonic_root
sendln "/usr/sbin/smartctl -x /dev/nvme0"
; ========== BMC version ==========
;wait prompt_sonic_root
;sendln "bmc-manager run 'cat /etc/issue'"
wait prompt_sonic_root
sendln "bmc-first-enroll"
wait prompt_sonic_root
sendln "bmc-manager version"
wait prompt_sonic_root
sendln "bmc-manager status"
; ========== Load Shell Script ==========
wait prompt_sonic_root
@@ -85,14 +222,13 @@ include "utils/show_pcie_error_reg_DDR.ttl"
include "utils/show_pcie_error_reg_I210.ttl"
include "utils/show_pcie_error_reg_SSD.ttl"
wait prompt_sonic_root
sendln "/usr/sbin/ras-mc-ctl --summary"
; ========== Check PCIE tree + PCIE link status + GET PCIE bandwidth Test ==========
include "utils/pcie_bus.ttl"
; ========== Check History ==========
;wait prompt_sonic_root
;sendln "command"
; ========== Check Margin ==========
if EN_Margin = 1 then
include "utils/show_margin_status.ttl"
@@ -101,33 +237,130 @@ endif
; ========== TAKE DATA ==========
include "utils/setup_pmon.ttl"
include "utils/show_dmesg.ttl"
wait prompt_sonic_root
;include "utils/xxx.ttl"
;wait prompt_sonic_root
;flushrecv
; ========== 10G/1G MGMT Ping Test ==========
wait prompt_sonic_root
sendln "./Blanton_Script/mgmt_ping_monitor.sh start"
wait prompt_sonic_root
pause 30
sendln "~/Blanton_Script/mgmt_ping_monitor.sh stop"
wait prompt_sonic_root
sendln "cat ~/Blanton_Script/log/mgmt_ping.log"
wait prompt_sonic_root
sendln "~/Blanton_Script/mgmt_ping_monitor.sh clear"
; ========== 100G Port Status ==========
wait prompt_sonic_root
sendln "sudo sfputil lpmode off Ethernet513"
wait prompt_sonic_root
sendln "sudo sfputil lpmode off Ethernet514"
wait prompt_sonic_root
sendln "config interface -n asic0 startup Ethernet513"
wait prompt_sonic_root
sendln "config interface -n asic1 startup Ethernet514"
wait prompt_sonic_root
sendln "config save -y"
wait prompt_sonic_root
sendln "show interfaces status Ethernet513,Ethernet514"
;=============================================================================================================
; ========== 100G PRBS ==========
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh start"
;wait prompt_sonic_root
;sendln "sleep 10"
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh status"
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh stop"
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh report"
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh clear"
; ========== 100G Port Status ==========
;wait prompt_sonic_root
;sendln "sudo sfputil lpmode off Ethernet513"
;wait prompt_sonic_root
;sendln "sudo sfputil lpmode off Ethernet514"
;wait prompt_sonic_root
;sendln "config interface -n asic0 startup Ethernet513"
;wait prompt_sonic_root
;sendln "config interface -n asic1 startup Ethernet514"
;wait prompt_sonic_root
;sendln "config save -y"
;wait prompt_sonic_root
;sendln "show interfaces status Ethernet513,Ethernet514"
;=============================================================================================================
; ========== First Traffic Test ==========
; Silent MAC
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps"
sendln "config feature state lldp disabled"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed stop"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report"
sendln "config save -y"
if SWB_UNIT0 = 1 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed stop"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report"
elseif SWB_UNIT0 = 1 && SWB_UNIT1 = 0 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear -u 0"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start -u 0"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed stop -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report -u 0"
elseif SWB_UNIT0 = 0 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear -u 1"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start -u 1"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed stop -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report -u 1"
else
endif
; ========== Show Power on uptime ==========
wait prompt_sonic_root
+119 -25
View File
@@ -20,63 +20,157 @@ sendln 'chmod +x ~/hammer/tools/amd/mlucas-avx2'
wait prompt_sonic_root
sendln "~/hammer/tools/amd/mlucas-avx2 -s l -cpu 0:15 > ~/hammer/tools/amd/mlucas_amm_log 2>&1 &"
; ========== BMC Stress Test ==========
; ========== BMC DDR Stress Test ==========
;wait prompt_sonic_root
;sendln "./Blanton_Script/bmc_monitor_ddr.sh start"
;wait prompt_sonic_root
;sendln "./Blanton_Script/bmc_monitor.sh start"
wait prompt_sonic_root
sendln "./Blanton_Script/bmc_monitor_ddr.sh start"
sendln "bgctl run bmc-manager run '/usr/bin/memtester 64M 1'"
; ========== BMC Monitor Test ==========
wait prompt_sonic_root
sendln "./Blanton_Script/bmc_monitor.sh start"
; ========== DDR MEMORY STRESS TEST ==========
;wait prompt_sonic_root
;sendln "bgctl run /usr/sbin/memtester 1G 100"
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_mem.py &"
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_mem.py &"
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_mem.py &"
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_mem.py &"
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_mem.py &"
sendln "bgctl run /usr/sbin/memtester 1G" ; Continue Execture
; ========== SSD read/write ==========
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_ssd.py &"
sendln "bgctl run qfx5252-stress-ssd 1G --runtime 3600 --passes 5 --write --force --log /host/hw-eval/jobs/qfx5252-stress-ssd.log"
; ========== USB read/write ==========
; ---- Get USB partition node -> usb_dev ----
usb_dev = ''
wait prompt_sonic_root
sendln "python ~/hammer/tools/stress_usb.py &"
sendln 'D=$(bash ~/Blanton_Script/usb_target.sh -o dev 2>/dev/null); echo "<<D""EV=${D:-NONE}>>"'
waitregex '<<DEV=([^>]*)>>'
usb_dev = groupmatchstr1
wait prompt_sonic_root
sprintf2 cmd "bgctl run qfx5252-stress-usb %s 256M --runtime 60 --passes 1 --write --log /host/hw-eval/jobs/qfx5252-stress-usb.log" usb_dev
sendln cmd
;wait prompt_sonic_root
;sendln "bgctl run qfx5252-stress-usb /dev/sda1 256M --runtime 60 --passes 1 --write --log /host/hw-eval/qfx5252-stress-usb.log"
; ========== SHOW Background job ==========
wait prompt_sonic_root
sendln "jobs"
wait prompt_sonic_root
sendln "bgctl list"
; ========== BMC USB Test ==========
wait prompt_sonic_root
sendln 'echo "NCMIF=$(ls -d /sys/bus/usb/drivers/cdc_ncm/*/net/* 2>/dev/null | head -n 1 | xargs -r basename)"'
timeout = 15
waitregex 'NCMIF=[A-Za-z0-9_-]+'
if result = 0 then
messagebox 'cdc_ncm interface not found. Ping test skipped.' 'ERROR'
goto skip_ping
endif
strlen matchstr
cut_len = result - 6 ; 'NCMIF=' = 6 chars
strcopy matchstr 7 cut_len ncm_if ; -> ncm_if = "eth2" / "eth3"
sprintf2 cmd 'ip -br link show %s' ncm_if
wait prompt_sonic_root
sendln cmd
wait prompt_sonic_root
sendln 'systemctl stop qfx5252-bmc-usb-net-test.service 2>/dev/null || true'
wait prompt_sonic_root
sendln 'systemctl reset-failed qfx5252-bmc-usb-net-test.service 2>/dev/null || true'
sprintf2 cmd 'systemd-run --unit=qfx5252-bmc-usb-net-test --property=Type=exec /usr/bin/ping -I %s -i 1 -W 2 -w 14400 192.168.200.200' ncm_if
wait prompt_sonic_root
sendln cmd
:skip_ping
; restore the default (no cap): the traffic init below walks 108 VLANs per
; unit and takes far longer than the 15 s set for the NCM probe above.
timeout = 0
; ========== 10G/1G MGMT Ping Test ==========
wait prompt_sonic_root
sendln "./Blanton_Script/mgmt_ping_monitor.sh start"
wait prompt_sonic_root
pause 3
sendln "./Blanton_Script/mgmt_ping_monitor.sh status"
; ========== 100G Port Status ==========
; ========== 100G Port ==========
wait prompt_sonic_root
sendln "show interfaces status Ethernet513,Ethernet514"
;=============================================================================================================
; ========== 100G PRBS ==========
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh start"
;=============================================================================================================
; ========== Traffic START ==========
; Silent MAC
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps"
sendln "config feature state lldp disabled"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start"
sendln "config save -y"
if SWB_UNIT0 = 1 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start"
elseif SWB_UNIT0 = 1 && SWB_UNIT1 = 0 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear -u 0"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start -u 0"
elseif SWB_UNIT0 = 0 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed ps -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed init -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed clear -u 1"
wait prompt_sonic_root
pause 15
sendln "blanton_traffic_linespeed show -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed start -u 1"
else
endif
; ========== Get data every 10mins ==========
; ========== Monitor loop: platform data + background jobs (pause 60s per round) ==========
while 1
include "utils/setup_pmon.ttl"
wait prompt_sonic_root
sendln "bgctl list"
wait prompt_sonic_root
sendln "jobs"
; ========== Check Margin ==========
if EN_Margin = 1 then
include "utils/show_margin_status.ttl"
endif
pause 60
endwhile
+119 -7
View File
@@ -12,6 +12,37 @@ pause 1
; ========== Kill Process ==========
include "utils/kill_all_process.ttl"
; bgctrl all stop
wait prompt_sonic_root
sendln "bgctl stop --all"
wait prompt_sonic_root
sendln "bgctl stop --all"
; BMC USB Test STOP
wait prompt_sonic_root
sendln "systemctl stop qfx5252-bmc-usb-net-test.service"
; 10G/1G MGMT Test STOP
wait prompt_sonic_root
sendln "~/Blanton_Script/mgmt_ping_monitor.sh status"
wait prompt_sonic_root
sendln "~/Blanton_Script/mgmt_ping_monitor.sh stop"
;=============================================================================================================
; 100G PRBS STOP
;wait prompt_sonic_root
;sendln "~/Blanton_Script/port_prbs_monitor.sh stop"
;wait prompt_sonic_root
;sendln "sudo sfputil lpmode off Ethernet513"
;wait prompt_sonic_root
;sendln "sudo sfputil lpmode off Ethernet514"
;wait prompt_sonic_root
;sendln "config interface -n asic0 startup Ethernet513"
;wait prompt_sonic_root
;sendln "config interface -n asic1 startup Ethernet514"
;=============================================================================================================
; ========== TAKE DATA ==========
include "utils/setup_pmon.ttl"
include "utils/show_dmesg.ttl"
@@ -25,28 +56,76 @@ include "utils/show_pcie_error_reg_DDR.ttl"
include "utils/show_pcie_error_reg_I210.ttl"
include "utils/show_pcie_error_reg_SSD.ttl"
; ========== BMC TAKE DATA ==========
wait prompt_sonic_root
sendln "./Blanton_Script/bmc_monitor_ddr.sh stop"
sendln "/usr/sbin/ras-mc-ctl --summary"
; ========== BMC DDR TAKE DATA ==========
;wait prompt_sonic_root
;sendln "./Blanton_Script/bmc_monitor_ddr.sh stop"
;wait prompt_sonic_root
;sendln "./Blanton_Script/bmc_monitor.sh stop"
;wait prompt_sonic_root
;sendln "cat ./Blanton_Script/log/bmc_poll.log"
;wait prompt_sonic_root
;sendln "cat ./Blanton_Script/log/bmc_ddr.log"
; ========== BMC Monitor TAKE DATA ==========
wait prompt_sonic_root
sendln "./Blanton_Script/bmc_monitor.sh stop"
wait prompt_sonic_root
sendln "cat ./Blanton_Script/log/bmc_poll.log"
wait prompt_sonic_root
sendln "cat ./Blanton_Script/log/bmc_ddr.log"
; ========== BMC USB Test result ==========
;wait prompt_sonic_root
;sendln "journalctl -u qfx5252-bmc-usb-net-test.service --no-pager"
; ========== CHECK Stress results ==========
wait prompt_sonic_root
sendln "cat ~/hammer/tools/amd/mlucas_amm_log"
wait prompt_sonic_root
sendln "cat ~/hammer/tools/log.stress_ssd"
sendln "cat /host/hw-eval/jobs/qfx5252-stress-ssd.log"
wait prompt_sonic_root
sendln "cat /host/hw-eval/jobs/qfx5252-stress-usb.log"
; ========== 10G/1G MGMT Test result ==========
wait prompt_sonic_root
sendln "cat ~/Blanton_Script/log/mgmt_ping.log"
; ========== 100G Port ==========
wait prompt_sonic_root
sendln "show interfaces status Ethernet513,Ethernet514"
;=============================================================================================================
;wait prompt_sonic_root
;sendln "cat ~/Blanton_Script/log/port_prbs.log "
;=============================================================================================================
; ========== CHECK Traffic counters ==========
if SWB_UNIT0 = 1 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed stop"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report"
elseif SWB_UNIT0 = 1 && SWB_UNIT1 = 0 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed stop -u 0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report -u 0"
elseif SWB_UNIT0 = 0 && SWB_UNIT1 = 1 then
wait prompt_sonic_root
sendln "blanton_traffic_linespeed stop -u 1"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report -u 1"
else
endif
; ========== NVME Err Info ==========
wait prompt_sonic_root
sendln "blanton_traffic_linespeed stop"
sendln "/usr/sbin/nvme smart-log -H /dev/nvme0"
wait prompt_sonic_root
sendln "blanton_traffic_linespeed report"
sendln "/usr/sbin/smartctl -x /dev/nvme0"
wait prompt_sonic_root
sendln "show reboot-cause"
@@ -55,4 +134,37 @@ sendln "show reboot-cause"
wait prompt_sonic_root
sendln "show uptime"
; ---- Get USB partition node -> usb_dev ----
usb_dev = ''
wait prompt_sonic_root
sendln 'D=$(bash ~/Blanton_Script/usb_target.sh -o dev 2>/dev/null); echo "<<D""EV=${D:-NONE}>>"'
waitregex '<<DEV=([^>]*)>>'
usb_dev = groupmatchstr1
wait prompt_sonic_root
sendln "sync"
wait prompt_sonic_root
sprintf2 cmd "mount %s /mnt/usb" usb_dev
sendln cmd
; ========== Copy Stress Log to USB ==========
wait prompt_sonic_root
sendln "cp ./Blanton_Script/log/bmc_poll.log /host/hw-eval/jobs/"
wait prompt_sonic_root
sendln "cp ./Blanton_Script/log/mgmt_ping.log /host/hw-eval/jobs/"
wait prompt_sonic_root
getdate ts_date "%Y%m%d"
gettime ts_time "%H%M"
sprintf2 usb_dst "/mnt/usb/jobs-%s_%s" ts_date ts_time
sprintf2 cmd "cp -r /host/hw-eval/jobs/ %s" usb_dst
sendln cmd
; ========== HW Test Session ==========
wait prompt_sonic_root
sendln "hw-test-session finish"
wait prompt_sonic_root
sendln "bgctl reset --yes"
messagebox 'GOOD JOB! Test Case DONE' 'teraterm'
@@ -14,6 +14,8 @@
| `margin_operation <ch> <op>` | 只切 OPERATION register |
| `margin_apply_profile <file>` | 批次套用 profile(套用在當前載入的 LTC2980 |
| `margin_save` | 寫入 NVM(當前 LTC2980 的兩個 LTC2977 都寫) |
| `margin_save all` | 對 `settings/` 底下**全部** LTC2980 寫入 NVM,結束後回到原本載入的那顆 |
| `margin_save blanton` | **Blanton SONiC 專用**18 顆 LTC2977 全部走原生 i2c bus 寫入,不需要 `margin_init` |
| `margin_debug [on\|off]` | 開關 debug 模式(不帶參數則 toggle |
| `margin_log [on\|off]` | 開關 log 記錄(不帶參數則 toggle,預設開啟) |
@@ -79,7 +81,9 @@ margin_init Blanton_SWB1_CONN16
margin_apply_profile profiles/comboA.conf
# 儲存到 NVM(斷電後保留)
margin_save
margin_save # 只寫當前載入的那顆
margin_save all # settings/ 底下 9 顆全部寫,最後自動切回原本那顆
margin_save blanton # Blanton SONiC 專用捷徑,18 顆一次寫完(見下)
```
## 一次掃所有 LTC2980
@@ -131,16 +135,87 @@ PROFILES=(
## 一次把所有 LTC2980 寫入 NVM
`script/margin_save_all.sh``BOARDS` 列出的每顆 LTC2980 跑 `margin_save`
> ⚠️ `margin_save` 會把當前 margin / OV / UV 設定**永久寫入 NVM**
> 斷電後保留。執行前先用 `margin_status_all.sh` 確認狀態正確。
**全部 9 顆**直接用內建的:
```bash
margin_save all
```
`settings/*.conf` 的檔名順序跑,中途某顆失敗**不會停**(半套完成的狀態比全部試完更難查),
最後印出「幾顆成功 / 哪幾顆失敗」的總結,並把 `margin_init` 切回你呼叫前載入的那顆。
**只寫其中幾顆**才需要 `script/margin_save_all.sh`,編 script 頂端的 `BOARDS` 陣列:
```bash
source script/margin_save_all.sh
```
要改寫哪幾顆,編 script 頂端的 `BOARDS` 陣列即可(同前述)。
### 送出去的是什麼
`margin_save` 對每顆 LTC2977 送一次 `STORE_USER_ALL`PMBus command `0x15`),
**不分 page**(這是整顆晶片的命令),也**不帶 data byte** —— 它是 SMBus send-byte
```bash
i2cset -y 4 0x5C 0x15 c # CB CONN13(原生 bus 4
cb_pmbus_write 8 0x5C 0x15 # SWB0 CONN13CB FPGA F3 Ch8
```
> v2.7.0 之前送的是 `0x15 0x00`write-byte)。LTC2977 在 CB 上吃得下去,但那是零件寬容,
> 不是規範。同時 v2.7.0 起會檢查回傳值 —— 以前 NACK 了還是印 `Change Saved`。
每顆之間會 `sleep $MARGIN_STORE_SETTLE`(預設 0.5 秒)。這不是量測值,是保守的保護:
NVM 寫入需要時間,而這支工具**沒有**去 poll 零件的 busy bit。`margin_save all` 連續送 18 次,
這個間隔就是唯一的緩衝。要關掉設成空字串即可。
## `margin_save blanton` — Blanton SONiC 專用
Blanton 這台把 CB FPGA F3 的 I2C 通道**同時掛成了原生 Linux i2c adapter**
所以 SWB 上的 LTC2977 不必走 `cb_pmbus_*` / pcimem,用 `i2cset` 直接打得到:
| FPGA F3 通道 | Linux bus | 板 / 接頭 |
|---|---|---|
| Ch6 | 13 | SWB0 CONN14 / CONN16 |
| Ch8 | 15 | SWB0 CONN13 / CONN15 |
| Ch7 | 14 | SWB1 CONN14 / CONN16 |
| Ch9 | 16 | SWB1 CONN13 / CONN15 |
| —(原生) | 4 | CB CONN13 |
因此 `margin_save blanton` 就是一張平表跑完 18 顆,**不需要 `margin_init`**、
不做 FPGA 通道初始化、不做 probe,也不會動到你當前載入的那顆板子:
```bash
margin_save blanton
```
```
CB CONN13 bus 4 0x5C : saved
SWB0 CONN13 bus 15 0x5C : saved
...
margin_save blanton: 18 saved, 0 failed
```
送出去的就是這 18 行(2026-08-27 上機驗證):
```bash
i2cset -f -y 4 0x5C 0x15 ; i2cset -f -y 4 0x5E 0x15 # CB CONN13
i2cset -f -y 15 0x5C 0x15 ; i2cset -f -y 15 0x5E 0x15 # SWB0 CONN13
i2cset -f -y 13 0x5C 0x15 ; i2cset -f -y 13 0x5E 0x15 # SWB0 CONN14
i2cset -f -y 15 0x62 0x15 ; i2cset -f -y 15 0x64 0x15 # SWB0 CONN15
i2cset -f -y 13 0x62 0x15 ; i2cset -f -y 13 0x64 0x15 # SWB0 CONN16
i2cset -f -y 16 0x5C 0x15 ; i2cset -f -y 16 0x5E 0x15 # SWB1 CONN13
i2cset -f -y 14 0x5C 0x15 ; i2cset -f -y 14 0x5E 0x15 # SWB1 CONN14
i2cset -f -y 16 0x62 0x15 ; i2cset -f -y 16 0x64 0x15 # SWB1 CONN15
i2cset -f -y 14 0x62 0x15 ; i2cset -f -y 14 0x64 0x15 # SWB1 CONN16
```
`-f` 是因為這些 bus 上的 LTC2977 可能已被 kernel driver 認走;上機驗證過的指令帶 `-f`,就照原樣保留。
**這台沒裝滿的話**,把 `margin.sh` 頂端 `MARGIN_BLANTON_STORE` 陣列裡對應的行註解掉即可,
不要讓它去打不存在的板子。bus 節點不存在(`/dev/i2c-N` 沒有)與晶片不回應(NACK
會分別報不同訊息 —— 前者代表 FPGA 的 i2c adapter 根本沒 enumerate,後者才是板子的問題。
## Operation 對應
@@ -424,8 +424,9 @@ SWB debug 指令(margin.sh v2.6.0 新增,皆為 `blanton_cb_i2c.sh` 的薄包裝
`Blanton_SWB0_CONN14.conf` 的第一顆也是 `0:0x5E`。兩個 CONN 現在同在 Ch6,
位址相撞。對照 SWB1 的規律(CONN13=`0x5C/0x5D`、CONN14=`0x5E/0x5F`),
SWB0_CONN13 第二顆推測應為 `0:0x5D`**尚未修改,待硬體確認**
- `margin_save``0x15` + data byte(`i2cset ... 0x15 0x00 b`)。PMBus 規範
STORE_USER_ALL 是 send-byte;CB 上實測可用故保留原樣,swb 路徑沿用同一寫法。
- `margin_save``0x15` **send-byte**(`i2cset ... 0x15 c` / `cb_pmbus_write <ch> <addr> 0x15`),
不帶 data byte。v2.7.0 之前送的是 `0x15 0x00`(write-byte),CB 上實測可用,
但那是零件寬容而非規範。
---
@@ -5,6 +5,24 @@
# =============================================================================
# Version Control
# -----------------------------------------------------------------------------
# v2.7.0 margin_save fixes and batching:
# * STORE_USER_ALL (0x15) is now issued as an SMBus send-byte instead
# of a write-byte with a dummy 0x00 data byte. The LTC2977 accepted
# the old form on CB, but it was the wrong protocol.
# i2c path: i2cset ... 0x15 c swb path: cb_pmbus_write ch addr 0x15
# * margin_save all -- STORE_USER_ALL on every board in settings/,
# then re-init whichever board was loaded beforehand. Does not stop
# on the first failure; prints a per-board summary at the end.
# * The store result is now checked. It used to be discarded, so a
# NACK still printed "Change Saved".
# * MARGIN_STORE_SETTLE (default 0.5s) between stores - nothing polls
# the part's busy bit, and `all` issues 18 of them in a row.
# * margin_save blanton -- Blanton SONiC shortcut. On this platform the
# CB FPGA F3 I2C channels also appear as native Linux i2c adapters
# (Ch6->bus 13, Ch7->14, Ch8->15, Ch9->16), so all 18 LTC2977 can be
# stored with plain i2cset: no margin_init, no FPGA channel setup,
# no pcimem. Bench-verified on the DUT 2026-08-27. The board list is
# MARGIN_BLANTON_STORE at the top of this file.
# v2.6.0 SWB path actually wired to the shared CB I2C package. blanton_cb_i2c.sh
# and blanton_fpga_pcimem.sh live in the PARENT dir (Blanton_Script/), not
# next to margin.sh, so the backend is now searched in both places.
@@ -115,6 +133,45 @@ ONOFF_ENABLE_VALUE="0x1a" # controlled_on=1, use_pmbus=1, use_control=0, b1=1
# margin_status after all channels are set. Bump this if you want margin_set's own
# read-back to reflect the settled value.
MARGIN_SETTLE="0.3"
# Pause after each STORE_USER_ALL. The LTC2977 is busy committing NVM once the
# command lands and nothing here polls a busy bit, so this is a plain guard
# rather than a measured value - it matters most for `margin_save all`, which
# fires 18 stores back to back. Set to "" to disable.
MARGIN_STORE_SETTLE="${MARGIN_STORE_SETTLE:-0.5}"
# --- `margin_save blanton`: the bench-verified Blanton SONiC store list -------
# On this platform the CB FPGA F3 I2C channels are ALSO exposed as native Linux
# i2c adapters, so every LTC2977 - CB and SWB alike - can be reached with plain
# i2cset. That path needs no FPGA channel setup, no pcimem, and no margin_init:
# one flat list, 18 commands, done. Verified on the DUT 2026-08-27.
#
# FPGA F3 channel -> Linux bus: Ch6->13 Ch7->14 Ch8->15 Ch9->16
#
# Entry format: "bus:addr:label". Comment out the lines for boards this DUT does
# not have rather than letting them fail - a single-SWB unit is a normal build.
MARGIN_BLANTON_STORE=(
"4:0x5C:CB CONN13"
"4:0x5E:CB CONN13"
"15:0x5C:SWB0 CONN13"
"15:0x5E:SWB0 CONN13"
"13:0x5C:SWB0 CONN14"
"13:0x5E:SWB0 CONN14"
"15:0x62:SWB0 CONN15"
"15:0x64:SWB0 CONN15"
"13:0x62:SWB0 CONN16"
"13:0x64:SWB0 CONN16"
"16:0x5C:SWB1 CONN13"
"16:0x5E:SWB1 CONN13"
"14:0x5C:SWB1 CONN14"
"14:0x5E:SWB1 CONN14"
"16:0x62:SWB1 CONN15"
"16:0x64:SWB1 CONN15"
"14:0x62:SWB1 CONN16"
"14:0x64:SWB1 CONN16"
)
# -f because the LTC2977s may be claimed by a kernel driver on these buses; the
# command that was verified on the bench carries it, so it stays.
MARGIN_BLANTON_I2CSET_FLAGS="${MARGIN_BLANTON_I2CSET_FLAGS:--f -y}"
_BUS=""
_ADDR=""
_PRE_SCOPE="" # set while inside a PRE/POST wrapped operation; prevents nested re-entry
@@ -257,6 +314,20 @@ _swb_write() {
return 0
}
# _swb_send <cmd> -> cb_pmbus_write with NO data byte (SMBus send-byte).
# cb_pmbus_write routes a no-data call to _cbi2c_xfer_cmd_only: START, slave+W,
# command byte, STOP. That is what STORE_USER_ALL expects.
_swb_send() {
local reg=$1
[ "$DEBUG_MODE" = "1" ] && echo -e "\033[90m[DEBG] cb_pmbus_write $CB_I2C_CH $_ADDR $reg (send-byte)\033[0m"
if ! _swb_call cb_pmbus_write "$CB_I2C_CH" "$_ADDR" "$reg" >/dev/null 2>&1; then
echo -e "[\033[31mERRO\033[0m] cb_pmbus_write ch=$CB_I2C_CH $_ADDR reg=$reg (send-byte) failed (NACK/timeout)" >&2
_log_detail "ERRO cb_pmbus_write ch=$CB_I2C_CH $_ADDR $reg send-byte"
return 1
fi
return 0
}
# _swb_read_word <reg> -> echoes "0xXXXX" parsed out of cb_pmbus_read's "=> 0xXXXX"
_swb_read_word() {
local reg=$1 out word
@@ -520,6 +591,20 @@ _i2c_write_byte() {
i2cset -y "$_BUS" "$_ADDR" "$1" "$2" b >/dev/null 2>&1
}
# Send a PMBus command with NO data byte (SMBus send-byte).
# STORE_USER_ALL (0x15) and its relatives are send-byte commands: the command
# byte alone triggers them. Appending a dummy 0x00 makes it a write-byte, which
# is a different SMBus protocol - the LTC2977 happened to accept it, but that is
# tolerance, not correctness. i2cset's "c" mode is the send-byte form.
_i2c_send_byte() {
if [ "$TRANSPORT" = "swb" ]; then
_swb_send "$1"
return
fi
[ "$DEBUG_MODE" = "1" ] && echo -e "\033[90m[DEBG] i2cset -y $_BUS $_ADDR $1 c\033[0m"
i2cset -y "$_BUS" "$_ADDR" "$1" c >/dev/null 2>&1
}
# Read a 16-bit word; echoes "0xXXXX". For swb, parse cb_pmbus_read's "=> 0xXXXX".
_i2c_read_word() {
if [ "$TRANSPORT" = "swb" ]; then
@@ -730,20 +815,171 @@ margin_apply_profile() {
}
# --- Save to NVM ---
# margin_save -- STORE_USER_ALL on the currently initialised board
# margin_save all -- ... on every board in settings/, then re-init the one
# that was loaded when you called it
# margin_save blanton -- Blanton SONiC shortcut: all 18 LTC2977 straight over
# the native i2c buses, no margin_init needed
#
# ⚠️ This is the only permanent write in the tool: it commits the current
# margin / OV / UV registers to NVM, and a power cycle will NOT undo it.
# Run margin_status first and never call it during a test run.
margin_save() {
if [ "${1:-}" = "all" ]; then
_margin_save_every_board
return
fi
if [ "${1:-}" = "blanton" ]; then
_margin_save_blanton
return
fi
local _own=0
_pre_enter_own && _own=1
local chip fails=0
for chip in "${CHIPS[@]}"; do
IFS=':' read -r _BUS _ADDR <<< "$chip"
_i2c_write_byte 0x15 0x00
# STORE_USER_ALL is chip-wide (all 8 pages), so no _open_page here.
if _i2c_send_byte 0x15; then
_log_detail "STORE_USER_ALL $PART_NUMBER $_ADDR ok"
else
fails=$(( fails + 1 ))
echo -e "[\033[31mERRO\033[0m] STORE_USER_ALL failed on $_ADDR" >&2
_log_detail "ERRO STORE_USER_ALL $PART_NUMBER $_ADDR"
fi
# The part is busy writing NVM after this; nothing here polls a busy bit,
# so leave a gap before touching it again. See MARGIN_STORE_SETTLE.
[ -n "$MARGIN_STORE_SETTLE" ] && sleep "$MARGIN_STORE_SETTLE"
done
echo "Change Saved (${#CHIPS[@]} chips)"
_log_label "margin_save $PART_NUMBER (${#CHIPS[@]} chips)"
if [ "$fails" -eq 0 ]; then
echo "Change Saved (${#CHIPS[@]} chips)"
else
echo -e "[\033[31mERRO\033[0m] $fails of ${#CHIPS[@]} chips did NOT save"
fi
_log_label "margin_save $PART_NUMBER (${#CHIPS[@]} chips, $fails failed)"
[ "$_own" = "1" ] && _pre_exit
return $(( fails > 0 ))
}
# margin_save all -- every .conf in settings/, in filename order.
# Deliberately does NOT stop on the first failure: a half-saved set is worse to
# diagnose than a fully-attempted one, and the summary tells you which boards
# to redo. margin_init is re-run per board because CHIPS / TRANSPORT /
# CB_I2C_CH are global state.
_margin_save_every_board() {
local files=( "$SETTINGS_DIR"/*.conf )
if [ ! -e "${files[0]}" ]; then
echo -e "[\033[31mERRO\033[0m] no .conf found in $SETTINGS_DIR"
return 1
fi
local restore="$PART_NUMBER" # empty if margin_init was never run
local board f ok=0 bad=0 failed=""
echo "-------------"
echo "margin_save all -- ${#files[@]} boards, permanent NVM write"
echo "-------------"
_log_label "margin_save all (${#files[@]} boards)"
for f in "${files[@]}"; do
board=$(basename "$f" .conf)
echo ""
if ! margin_init "$board"; then
bad=$(( bad + 1 )); failed="$failed $board(init)"
continue
fi
if margin_save; then
ok=$(( ok + 1 ))
else
bad=$(( bad + 1 )); failed="$failed $board"
fi
done
echo ""
echo "-------------"
printf "margin_save all: %d saved, %d failed\n" "$ok" "$bad"
[ "$bad" -gt 0 ] && echo -e "[\033[31mERRO\033[0m] failed:$failed"
echo "-------------"
_log_label "margin_save all done: $ok ok, $bad failed$failed"
# Put back whatever board the caller had loaded, so a following
# margin_status does not silently report the last board in the list.
if [ -n "$restore" ] && [ "$restore" != "$board" ]; then
echo ""
echo -e "[\033[34mINFO\033[0m] restoring previous board: $restore"
margin_init "$restore" >/dev/null
fi
return $(( bad > 0 ))
}
# margin_save blanton -- Blanton SONiC only. Walks MARGIN_BLANTON_STORE and
# sends STORE_USER_ALL to each LTC2977 over the native i2c buses.
#
# Deliberately does NOT use margin_init / CHIPS / TRANSPORT: the whole point is
# that on this platform every LTC2977 is reachable with plain i2cset, so the
# FPGA channel setup and the per-board probe are not needed. That also means it
# does not touch the loaded board - PART_NUMBER is the same afterwards.
_margin_save_blanton() {
local n=${#MARGIN_BLANTON_STORE[@]}
if [ "$n" -eq 0 ]; then
echo -e "[\033[31mERRO\033[0m] MARGIN_BLANTON_STORE is empty"
return 1
fi
echo "-------------"
echo "margin_save blanton -- $n LTC2977, permanent NVM write"
echo "-------------"
_log_label "margin_save blanton ($n chips)"
local entry bus addr label ok=0 bad=0 failed=""
for entry in "${MARGIN_BLANTON_STORE[@]}"; do
IFS=':' read -r bus addr label <<< "$entry"
# "no such bus" and "chip did not answer" are different faults: the first
# means the FPGA i2c adapters did not enumerate, the second means the
# board is absent or the address is wrong. Do not blur them into one.
if [ ! -e "/dev/i2c-$bus" ]; then
echo -e "[\033[31mERRO\033[0m] $label bus $bus addr $addr : /dev/i2c-$bus does not exist"
_log_detail "ERRO STORE_USER_ALL $label bus=$bus addr=$addr no-such-bus"
bad=$(( bad + 1 )); failed="$failed [$label $addr]"
continue
fi
[ "$DEBUG_MODE" = "1" ] && echo -e "\033[90m[DEBG] i2cset $MARGIN_BLANTON_I2CSET_FLAGS $bus $addr 0x15\033[0m"
if i2cset $MARGIN_BLANTON_I2CSET_FLAGS "$bus" "$addr" 0x15 >/dev/null 2>&1; then
printf " %-12s bus %-2s %s : saved\n" "$label" "$bus" "$addr"
ok=$(( ok + 1 ))
_log_detail "STORE_USER_ALL $label bus=$bus addr=$addr ok"
else
echo -e "[\033[31mERRO\033[0m] $label bus $bus addr $addr : i2cset failed (NACK/busy)"
_log_detail "ERRO STORE_USER_ALL $label bus=$bus addr=$addr"
bad=$(( bad + 1 )); failed="$failed [$label $addr]"
fi
# No busy-bit poll anywhere in this tool; this gap is the only buffer.
[ -n "$MARGIN_STORE_SETTLE" ] && sleep "$MARGIN_STORE_SETTLE"
done
echo "-------------"
printf "margin_save blanton: %d saved, %d failed\n" "$ok" "$bad"
[ "$bad" -gt 0 ] && echo -e "[\033[31mERRO\033[0m] failed:$failed"
echo "-------------"
_log_label "margin_save blanton done: $ok ok, $bad failed$failed"
return $(( bad > 0 ))
}
# --- Tab completion for margin_save ---
_margin_save_completions() {
[ "$COMP_CWORD" -ne 1 ] && return
COMPREPLY=($(compgen -W "all blanton" -- "${COMP_WORDS[COMP_CWORD]}"))
}
complete -F _margin_save_completions margin_save
if [ -z "$_LOG_FILE" ]; then
_LOG_FILE="$LOGS_DIR/margin_$(date '+%Y%m%d_%H%M%S').log"
_LOG_DETAIL_FILE="${_LOG_FILE%.log}_detail.log"
@@ -4,6 +4,10 @@
# ⚠️ margin_save 會把目前的 margin / OV / UV 設定**永久寫入 NVM**
# 斷電後仍保留。執行前請務必先 margin_status 確認各 channel 狀態正確。
#
# 全部 9 顆請直接用 margin.sh v2.7.0 內建的 `margin_save all` —— 它有成功/失敗總結,
# 跑完還會把 margin_init 切回你原本載入的那顆。這支只在「只想寫其中幾顆」時才需要,
# 編下面的 BOARDS 陣列。
#
# Usage:
# source script/margin_save_all.sh
@@ -34,6 +34,10 @@ set -u
# Add a line to add a command; comment it out to disable it.
COMMANDS=(
"free -m"
"i2ctransfer -f -y 1 w6@0x41 0x00 0xC0 0x55 0xAA 0x55 0xAA"
"i2ctransfer -f -y 1 w2@0x41 0x00 0xC0 r4"
"i2ctransfer -f -y 1 w6@0x41 0x00 0xC0 0xAA 0x55 0xAA 0x55"
"i2ctransfer -f -y 1 w2@0x41 0x00 0xC0 r4"
#"uptime"
#"cat /proc/loadavg"
#"cat /proc/meminfo"
@@ -0,0 +1,403 @@
#!/bin/bash
###############################################################################
# mgmt_ping_monitor.sh
#
# Version : V3.0.0
# Author : ETWen
# Date : 20260826
#
# Purpose : Ping from BOTH management NICs at the same time, continuously, and
# keep accumulating until stopped:
#
# 10G (eth0) -> TARGET_10G
# 1G (eth1) -> TARGET_1G
#
# `stop` renders a report into the log: the last TAIL_LINES entries
# per NIC, each one's ping statistics and verdict, then
# `ip -s link show` for both interfaces.
#
# Notes : - Both NICs are configured with iproute2 `replace`, which is
# idempotent, and both are left up. On a shared subnet the target's
# ARP can be answered by either NIC, so ARP_STRICT keeps each one
# to its own address.
# - ARP WARM-UP: the first packet after a neighbour entry expires is
# spent resolving ARP and is counted as loss. On the bench this
# produced three FAILs whose only missing packet was icmp_seq=1,
# every time, with zero NIC errors or drops. One discarded ping
# before the measured run removes that artefact, which is what lets
# MAX_LOSS_PCT stay at 0 and still mean something.
# - MANAGEMENT CONNECTIVITY IS IN USE while this runs. Drive it from
# the serial console.
#
# Version History
# V1.0.0 20260824 Initial Version (ifconfig, one NIC at a time)
# V2.0.0 20260824 iproute2; both NICs stay up; per-NIC target;
# ip -s link show + TX/RX delta per leg
# V3.0.0 20260826 Both NICs ping SIMULTANEOUSLY and continuously instead
# of alternating fixed bursts. Add the ARP warm-up.
# `stop` renders last-N per NIC + statistics + both
# ip -s link outputs into the log.
###############################################################################
set -u
###############################################################################
# User Configurable Section -- normally this is the only part you need to edit
###############################################################################
# --- 10G management port ---
IF_10G="eth0"
IP_10G="192.168.1.99"
TARGET_10G="192.168.1.30" # host this NIC pings
PLEN_10G=24 # prefix length, e.g. 24 = /24
GW_10G="" # default gateway; empty = do not touch routing
METRIC_10G=100 # lower metric wins for off-subnet traffic
# --- 1G management port ---
IF_1G="eth1"
IP_1G="192.168.1.101"
TARGET_1G="192.168.1.31"
PLEN_1G=24
GW_1G=""
METRIC_1G=200
# --- ping ---
PING_INTERVAL_SEC=1 # seconds between echo requests (ping -i)
MAX_LOSS_PCT=0 # loss above this marks the NIC FAIL
TAIL_LINES=20 # how many recent entries per NIC the report shows
# Extra ping flags. -D timestamps every line, -O prints a marker for a request
# that got no reply, so a drop is visible in the tail instead of just missing.
# Clear this if the platform's ping does not accept them.
PING_EXTRA_OPTS="-D -O"
# One discarded ping per NIC before the measured run, to resolve ARP. Without
# it the first packet of the run is lost to neighbour resolution and looks
# exactly like a link fault. 0 disables.
ARP_WARMUP=1
ARP_WARMUP_TIMEOUT=2 # seconds to wait for the warm-up reply
# Both NICs stay up. If they share a subnet the target's ARP can be answered by
# either one, so a reply may arrive on the NIC that did not send. This applies
# arp_ignore=1 / arp_announce=2 to both. RAM only -- reverts on reboot.
ARP_STRICT=1
# Prefix for the ip commands; empty because this already runs as root.
SUDO=""
# Seconds to wait after bringing a link up before pinging.
LINK_SETTLE_SEC=5
# Logging
LOG_DIR="" # empty = <script directory>/log
LOG_NAME="mgmt_ping.log"
# What `start` does with the log left behind by the previous run:
# new = discard it (default -- one run, one log)
# archive = rename it to <log>.YYYYmmdd-HHMMSS first
# append = keep it
START_LOG_MODE="new"
###############################################################################
# Internal
###############################################################################
SCRIPT_NAME="$(basename -- "${BASH_SOURCE[0]}")"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
[ -n "${LOG_DIR}" ] || LOG_DIR="${SCRIPT_DIR}/log"
LOG_FILE="${LOG_DIR}/${LOG_NAME}"
BASE="${LOG_DIR}/${LOG_NAME%.log}"
PID_FILE="${BASE}.pid"
STATE_FILE="${BASE}.state"
RAW_10G="${BASE}_${IF_10G}.raw"
RAW_1G="${BASE}_${IF_1G}.raw"
ts() { date '+%Y-%m-%d %H:%M:%S'; }
die() { printf '[ERROR] %s\n' "$*" >&2; exit 1; }
# "<rx_packets> <tx_packets>" for an interface. ip -s link orders the columns
# "bytes packets errors ...", so packets is $2, not $1.
if_counters() {
${SUDO} ip -s link show "$1" 2>/dev/null | awk '
/RX:/ { getline; rx = $2 }
/TX:/ { getline; tx = $2 }
END { printf "%s %s", (rx == "" ? 0 : rx), (tx == "" ? 0 : tx) }'
}
# Bring the NIC up and (re)apply its address and default route.
setup_iface() {
local iface="$1" ipaddr="$2" plen="$3" gw="$4" metric="$5"
${SUDO} ip link set "${iface}" up 2>&1 || return 1
${SUDO} ip address replace "${ipaddr}/${plen}" dev "${iface}" 2>&1 || return 1
if [ -n "${gw}" ]; then
${SUDO} ip route replace default via "${gw}" dev "${iface}" \
metric "${metric}" 2>&1 || return 1
fi
return 0
}
apply_arp_strict() {
local i
[ "${ARP_STRICT}" -eq 1 ] || return 0
for i in "${IF_10G}" "${IF_1G}"; do
${SUDO} sysctl -qw "net.ipv4.conf.${i}.arp_ignore=1" 2>/dev/null
${SUDO} sysctl -qw "net.ipv4.conf.${i}.arp_announce=2" 2>/dev/null
done
}
# Resolve the neighbour so the measured run does not spend its first packet on
# ARP. Result is deliberately discarded.
arp_warmup() {
local iface="$1" target="$2"
[ "${ARP_WARMUP}" -eq 1 ] || return 0
ping -I "${iface}" -c 1 -W "${ARP_WARMUP_TIMEOUT}" "${target}" >/dev/null 2>&1
return 0
}
# echo "pid_10g pid_1g" and return 0 when both are alive
is_running() {
local p0 p1
[ -f "${PID_FILE}" ] || return 1
p0="$(sed -n '1p' "${PID_FILE}" 2>/dev/null)"
p1="$(sed -n '2p' "${PID_FILE}" 2>/dev/null)"
[[ "${p0}" =~ ^[0-9]+$ ]] || return 1
[[ "${p1}" =~ ^[0-9]+$ ]] || return 1
kill -0 "${p0}" 2>/dev/null || kill -0 "${p1}" 2>/dev/null || return 1
printf '%s %s' "${p0}" "${p1}"
return 0
}
###############################################################################
# Report
###############################################################################
# The recent entries for one NIC: replies and, thanks to -O, the requests that
# got none.
tail_entries() {
grep -aE 'bytes from|no answer|Unreachable|Time to live' "$1" 2>/dev/null \
| tail -n "${TAIL_LINES}"
}
# The trailing "--- x ping statistics ---" block.
stats_block() {
sed -n '/ping statistics ---/,$p' "$1" 2>/dev/null
}
# render_leg <label> <iface> <target> <raw> <rx0> <tx0>
render_leg() {
local label="$1" iface="$2" target="$3" raw="$4" rx0="$5" tx0="$6"
local st tx rx loss avg verdict c1 rx1 tx1 drx dtx
printf -- '----- %s : %s -> %s : last %s entries -----\n' \
"${label}" "${iface}" "${target}" "${TAIL_LINES}"
tail_entries "${raw}"
printf '\n'
st="$(stats_block "${raw}")"
printf '%s\n' "${st}"
tx="$(printf '%s' "${st}" | sed -n 's/^\([0-9]\+\) packets transmitted.*/\1/p' | tail -1)"
rx="$(printf '%s' "${st}" | sed -n 's/.*[, ]\([0-9]\+\) received.*/\1/p' | tail -1)"
loss="$(printf '%s' "${st}" | sed -n 's/.*[, ]\([0-9]\+\)% packet loss.*/\1/p' | tail -1)"
avg="$(printf '%s' "${st}" | sed -n 's|.*= [0-9.]*/\([0-9.]*\)/.*|\1|p' | tail -1)"
[ -n "${tx}" ] || tx=0
[ -n "${rx}" ] || rx=0
[ -n "${loss}" ] || loss=100
[ -n "${avg}" ] || avg="-"
c1="$(if_counters "${iface}")"; rx1="${c1% *}"; tx1="${c1#* }"
drx=$(( rx1 - rx0 )); dtx=$(( tx1 - tx0 ))
if [ "${loss}" -le "${MAX_LOSS_PCT}" ] && [ "${tx}" -gt 0 ]; then
verdict="PASS"
else
verdict="FAIL"
fi
# nic_tx/nic_rx are this interface's own counter delta over the whole run.
# A pass with nic_tx near zero means the traffic left on the other NIC.
printf '[%s] RESULT %-3s %-6s -> %-15s tx=%s rx=%s loss=%s%% rtt_avg=%sms nic_tx=%s nic_rx=%s %s\n\n' \
"$(ts)" "${label}" "${iface}" "${target}" "${tx}" "${rx}" "${loss}" "${avg}" \
"${dtx}" "${drx}" "${verdict}"
}
write_report() {
local rx0_10g tx0_10g rx0_1g tx0_1g started
started="$(sed -n '1p' "${STATE_FILE}" 2>/dev/null)"
rx0_10g="$(sed -n '2p' "${STATE_FILE}" 2>/dev/null)"; rx0_10g="${rx0_10g:-0}"
tx0_10g="$(sed -n '3p' "${STATE_FILE}" 2>/dev/null)"; tx0_10g="${tx0_10g:-0}"
rx0_1g="$(sed -n '4p' "${STATE_FILE}" 2>/dev/null)"; rx0_1g="${rx0_1g:-0}"
tx0_1g="$(sed -n '5p' "${STATE_FILE}" 2>/dev/null)"; tx0_1g="${tx0_1g:-0}"
{
printf '#############################################################\n'
printf '[%s] mgmt ping report\n' "$(ts)"
printf ' started : %s\n' "${started:-unknown}"
printf ' stopped : %s\n' "$(ts)"
printf ' 10G : %s %s/%s -> %s (metric %s)\n' \
"${IF_10G}" "${IP_10G}" "${PLEN_10G}" "${TARGET_10G}" "${METRIC_10G}"
printf ' 1G : %s %s/%s -> %s (metric %s)\n' \
"${IF_1G}" "${IP_1G}" "${PLEN_1G}" "${TARGET_1G}" "${METRIC_1G}"
printf ' ping : -i %s %s, max loss %s%%\n' \
"${PING_INTERVAL_SEC}" "${PING_EXTRA_OPTS}" "${MAX_LOSS_PCT}"
printf '#############################################################\n\n'
render_leg "10G" "${IF_10G}" "${TARGET_10G}" "${RAW_10G}" "${rx0_10g}" "${tx0_10g}"
render_leg "1G" "${IF_1G}" "${TARGET_1G}" "${RAW_1G}" "${rx0_1g}" "${tx0_1g}"
printf -- '--- ip -s link show %s ---\n' "${IF_10G}"
${SUDO} ip -s link show "${IF_10G}" 2>&1
printf '\n'
printf -- '--- ip -s link show %s ---\n' "${IF_1G}"
${SUDO} ip -s link show "${IF_1G}" 2>&1
} >> "${LOG_FILE}"
}
###############################################################################
# Sub-commands
###############################################################################
prepare_log_on_start() {
local stamp
case "${START_LOG_MODE}" in
append) return 0 ;;
archive)
if [ -s "${LOG_FILE}" ]; then
stamp="$(date '+%Y%m%d-%H%M%S')"
mv -f "${LOG_FILE}" "${LOG_FILE}.${stamp}" || die "cannot archive ${LOG_FILE}"
printf 'previous log archived: %s.%s\n' "${LOG_FILE}" "${stamp}"
fi
;;
new)
# Truncate rather than unlink, so a tail -f already attached keeps
# following the new run.
[ -f "${LOG_FILE}" ] && { : > "${LOG_FILE}" || die "cannot truncate ${LOG_FILE}"; }
;;
*) die "invalid START_LOG_MODE: ${START_LOG_MODE}" ;;
esac
}
do_start() {
local pids p0 p1 c
pids="$(is_running)" && die "already running (pids=${pids})"
command -v ip >/dev/null 2>&1 || die "ip (iproute2) not found"
command -v ping >/dev/null 2>&1 || die "ping not found"
mkdir -p "${LOG_DIR}" || die "cannot create ${LOG_DIR}"
prepare_log_on_start
: > "${RAW_10G}"
: > "${RAW_1G}"
setup_iface "${IF_10G}" "${IP_10G}" "${PLEN_10G}" "${GW_10G}" "${METRIC_10G}" \
|| die "cannot configure ${IF_10G}"
setup_iface "${IF_1G}" "${IP_1G}" "${PLEN_1G}" "${GW_1G}" "${METRIC_1G}" \
|| die "cannot configure ${IF_1G}"
apply_arp_strict
sleep "${LINK_SETTLE_SEC}"
# Spend the ARP resolution here, not on the first measured packet.
arp_warmup "${IF_10G}" "${TARGET_10G}"
arp_warmup "${IF_1G}" "${TARGET_1G}"
# Counter baseline, taken after the warm-up so its packets are excluded.
{ c="$(if_counters "${IF_10G}")"
printf '%s\n%s\n%s\n' "$(ts)" "${c% *}" "${c#* }"
c="$(if_counters "${IF_1G}")"
printf '%s\n%s\n' "${c% *}" "${c#* }"
} > "${STATE_FILE}"
# Both NICs ping at the same time and keep accumulating until stopped.
nohup ping -I "${IF_10G}" -i "${PING_INTERVAL_SEC}" ${PING_EXTRA_OPTS} \
"${TARGET_10G}" >> "${RAW_10G}" 2>&1 &
p0=$!
nohup ping -I "${IF_1G}" -i "${PING_INTERVAL_SEC}" ${PING_EXTRA_OPTS} \
"${TARGET_1G}" >> "${RAW_1G}" 2>&1 &
p1=$!
disown "${p0}" 2>/dev/null
disown "${p1}" 2>/dev/null
printf '%s\n%s\n' "${p0}" "${p1}" > "${PID_FILE}"
sleep 1
kill -0 "${p0}" 2>/dev/null || die "10G ping failed to start, see ${RAW_10G}"
kill -0 "${p1}" 2>/dev/null || die "1G ping failed to start, see ${RAW_1G}"
printf 'started (10G pid=%s, 1G pid=%s)\nlog: %s\n' "${p0}" "${p1}" "${LOG_FILE}"
}
do_stop() {
local pids p0 p1 i
pids="$(is_running)" || {
printf 'not running\n'
rm -f "${PID_FILE}"
return 0
}
p0="${pids% *}"; p1="${pids#* }"
# SIGINT, not SIGTERM: ping prints its statistics block on interrupt, and
# that block is what the report parses.
kill -INT "${p0}" 2>/dev/null
kill -INT "${p1}" 2>/dev/null
for (( i = 0; i < 20; i++ )); do
kill -0 "${p0}" 2>/dev/null || kill -0 "${p1}" 2>/dev/null || break
sleep 0.5
done
kill -KILL "${p0}" 2>/dev/null
kill -KILL "${p1}" 2>/dev/null
write_report
rm -f "${PID_FILE}"
printf 'stopped (10G pid=%s, 1G pid=%s)\nreport: %s\n' "${p0}" "${p1}" "${LOG_FILE}"
}
do_status() {
local pids
if pids="$(is_running)"; then
printf 'status : running (10G pid=%s, 1G pid=%s)\n' "${pids% *}" "${pids#* }"
else
printf 'status : stopped\n'
fi
printf 'log : %s\n' "${LOG_FILE}"
printf '10G : %s %s/%s -> %s\n' "${IF_10G}" "${IP_10G}" "${PLEN_10G}" "${TARGET_10G}"
printf '1G : %s %s/%s -> %s\n' "${IF_1G}" "${IP_1G}" "${PLEN_1G}" "${TARGET_1G}"
[ -f "${RAW_10G}" ] && printf '10G recv: %s\n' "$(grep -ac 'bytes from' "${RAW_10G}")"
[ -f "${RAW_1G}" ] && printf '1G recv: %s\n' "$(grep -ac 'bytes from' "${RAW_1G}")"
}
usage() {
cat <<EOF
Usage: ${SCRIPT_NAME} {start|stop|status|summary|clear}
start Configure both NICs, warm up ARP, then ping from BOTH at the same
time and keep accumulating. The previous log is discarded first.
stop Stop both pings and render the report into the log
status Show pids, configured NICs and replies received so far
summary Print just the RESULT lines from the log
clear Remove the log and the raw captures (must be stopped first)
Log file : ${LOG_FILE}
The report holds, per NIC, the last ${TAIL_LINES} entries and the ping
statistics, then "ip -s link show" for both interfaces.
Both NICs are used at once, so management connectivity is in play -- drive
this from the serial console. Edit the User Configurable Section at the top
to change interfaces, addresses, targets or the ping options.
EOF
}
###############################################################################
# Entry point
###############################################################################
case "${1:-}" in
start) do_start ;;
stop) do_stop ;;
status) do_status ;;
summary) grep -a 'RESULT' "${LOG_FILE}" 2>/dev/null || printf 'no results in %s\n' "${LOG_FILE}" ;;
clear)
is_running >/dev/null && die "still running, stop it first"
rm -f "${LOG_FILE}" "${LOG_FILE}".[0-9]* "${RAW_10G}" "${RAW_1G}" "${STATE_FILE}"
printf 'log cleared\n'
;;
-h|--help|"") usage ;;
*) printf '[ERROR] unknown sub-command: %s\n\n' "$1" >&2; usage; exit 1 ;;
esac
@@ -0,0 +1,431 @@
#!/bin/bash
###############################################################################
# port_prbs_monitor.sh
#
# Version : V1.0.0
# Author : ETWen
# Date : 20260826
#
# Purpose : Run a PRBS test on the two 100G uplinks AT THE SAME TIME and keep
# polling until stopped. Each port gets its own background worker:
#
# setup : sfputil lpmode off <port>
# phy diag <phy> prbs set <poly>
# phy diag <phy> prbsstat STArt Interval=<n>
# poll : phy diag <phy> prbs get <- PASS/FAIL comes from here
# phydiag <phy> prbsstat Ber <- recorded only
# stop : phydiag <phy> prbsstat STOp
# phydiag <phy> prbs clear
#
# Every command and its full output goes to that port's raw log.
# `stop` tears the test down and renders the report; `report`
# prints the PASS/FAIL tally per port.
#
# Notes : - A poll counts as PASS only when the output contains "PRBS OK!".
# bcmcmd exits 0 even when the BCM shell rejects a command, so the
# exit status cannot be used -- only the output can.
# - bcmcmd is always run with </dev/null. Without it, it inherits the
# polling loop's stdin and eats it, and the loop runs once.
#
# Version History
# V1.0.0 20260826 Initial Version
###############################################################################
set -u
###############################################################################
# User Configurable Section -- normally this is the only part you need to edit
###############################################################################
# --- port A ---
PORT_A_NAME="Ethernet513"
PORT_A_UNIT=0 # bcmcmd -n <unit>
PORT_A_PHY=268 # phy diag <phy>
# --- port B ---
PORT_B_NAME="Ethernet514"
PORT_B_UNIT=1
PORT_B_PHY=268
# Which ports `start` runs when no argument is given: both | a | b
# Override per run: start a / start b / start both
# Running one at a time is how you tell a genuine port fault from the DUT not
# coping with two PRBS streams at once.
PORTS="both"
# --- PRBS ---
PRBS_POLY="p=3" # passed to "prbs set"
PRBS_STAT_INTERVAL=5 # passed to "prbsstat STArt Interval="
POLL_INTERVAL_SEC=10 # seconds between polls
PASS_PATTERN="PRBS OK!" # a poll is PASS only if the output contains this
# Seconds to wait after the setup commands before the first poll, so the link
# has settled and the counters mean something.
SETTLE_SEC=10
# Stop after this many polls per port. 0 = run until stopped manually.
MAX_POLLS=0
# How many recent poll blocks the report shows per port.
TAIL_ENTRIES=20
# Prefix for privileged commands; empty because this already runs as root.
SUDO=""
BCMCMD="bcmcmd"
# Logging
LOG_DIR="" # empty = <script directory>/log
LOG_NAME="port_prbs.log"
# What `start` does with the log left behind by the previous run:
# new | archive | append
START_LOG_MODE="new"
###############################################################################
# Internal
###############################################################################
SCRIPT_NAME="$(basename -- "${BASH_SOURCE[0]}")"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
SCRIPT_PATH="${SCRIPT_DIR}/${SCRIPT_NAME}"
[ -n "${LOG_DIR}" ] || LOG_DIR="${SCRIPT_DIR}/log"
LOG_FILE="${LOG_DIR}/${LOG_NAME}"
BASE="${LOG_DIR}/${LOG_NAME%.log}"
PID_FILE="${BASE}.pid"
STATE_FILE="${BASE}.state"
RAW_A="${BASE}_${PORT_A_NAME}.raw"
RAW_B="${BASE}_${PORT_B_NAME}.raw"
ts() { date '+%Y-%m-%d %H:%M:%S'; }
# port_tally <name> <raw> -> "<polls> <pass> <fail> <verdict>"
# A raw without a "PRBS START" line means the port was not run this session,
# which is SKIP -- reporting it as FAIL would make a deliberate single-port
# run look like half the hardware is broken.
port_tally() {
local name="$1" raw="$2" pass fail polls verdict
if [ ! -s "${raw}" ] || ! grep -aq "PRBS START ${name}" "${raw}" 2>/dev/null; then
printf '0 0 0 SKIP'; return 0
fi
pass=$(grep -ac "PRBS ${name} poll=[0-9]* PASS" "${raw}" 2>/dev/null); pass=${pass:-0}
fail=$(grep -ac "PRBS ${name} poll=[0-9]* FAIL" "${raw}" 2>/dev/null); fail=${fail:-0}
polls=$(( pass + fail ))
if [ "${polls}" -gt 0 ] && [ "${fail}" -eq 0 ]; then verdict=PASS; else verdict=FAIL; fi
printf '%s %s %s %s' "${polls}" "${pass}" "${fail}" "${verdict}"
}
die() { printf '[ERROR] %s\n' "$*" >&2; exit 1; }
# bcm <unit> <dsh command> -- output on stdout, never trusts the exit status.
# </dev/null so it cannot consume the caller's stdin.
bcm() {
local unit="$1" cmd="$2"
${SUDO} "${BCMCMD}" -n "${unit}" -c "dsh -c \"${cmd}\"" </dev/null 2>&1
}
# run_step <unit> <dsh command>
# The command and its full output are written to the LOG via stderr, because
# the worker has stderr pointed at the raw file. Only the output itself goes to
# stdout, so `out="$(run_step ...)"` captures the output without swallowing the
# log lines -- writing both to stdout would have put the whole record inside
# the variable and left the log with nothing but the poll verdicts.
run_step() {
local unit="$1" cmd="$2" out
out="$(bcm "${unit}" "${cmd}")"
{
printf '[%s] CMD : bcmcmd -n %s -c '\''dsh -c "%s"'\''\n' "$(ts)" "${unit}" "${cmd}"
printf '%s\n' "${out}"
} >&2
printf '%s' "${out}"
}
# Echo "<pid_a> <pid_b>" and return 0 while at least one worker is alive. A
# port that was not started this run is recorded as "-", so a single-port run
# is a first-class case rather than a half-broken two-port one.
is_running() {
local pa pb alive=0
[ -f "${PID_FILE}" ] || return 1
pa="$(sed -n '1p' "${PID_FILE}" 2>/dev/null)"; pa="${pa:--}"
pb="$(sed -n '2p' "${PID_FILE}" 2>/dev/null)"; pb="${pb:--}"
[[ "${pa}" =~ ^[0-9]+$ ]] && kill -0 "${pa}" 2>/dev/null && alive=1
[[ "${pb}" =~ ^[0-9]+$ ]] && kill -0 "${pb}" 2>/dev/null && alive=1
[ "${alive}" -eq 1 ] || return 1
printf '%s %s' "${pa}" "${pb}"
return 0
}
###############################################################################
# Worker -- one per port, launched by `start`
###############################################################################
WORKER_STOP=0
worker_on_signal() { WORKER_STOP=1; }
# __worker <name> <unit> <phy> <rawfile>
do_worker() {
local name="$1" unit="$2" phy="$3" raw="$4"
local n=0 pass=0 fail=0 out
exec >>"${raw}" 2>&1
trap worker_on_signal INT TERM
printf '#############################################################\n'
printf '[%s] PRBS START %s (unit %s, phy %s) poly=%s interval=%ss poll=%ss\n' \
"$(ts)" "${name}" "${unit}" "${phy}" "${PRBS_POLY}" \
"${PRBS_STAT_INTERVAL}" "${POLL_INTERVAL_SEC}"
printf '#############################################################\n'
# --- setup ---
printf '[%s] CMD : sfputil lpmode off %s\n' "$(ts)" "${name}"
${SUDO} sfputil lpmode off "${name}" </dev/null 2>&1
run_step "${unit}" "phy diag ${phy} prbs set ${PRBS_POLY}" >/dev/null
run_step "${unit}" "phy diag ${phy} prbsstat STArt Interval=${PRBS_STAT_INTERVAL}" >/dev/null
sleep "${SETTLE_SEC}"
# --- poll ---
while [ "${WORKER_STOP}" -eq 0 ]; do
n=$(( n + 1 ))
printf -- '---------- %s POLL %d @ %s ----------\n' "${name}" "${n}" "$(ts)"
out="$(run_step "${unit}" "phy diag ${phy} prbs get")"
# Ber is recorded for the log only; it does not decide the verdict.
run_step "${unit}" "phydiag ${phy} prbsstat Ber" >/dev/null
if printf '%s' "${out}" | grep -qF "${PASS_PATTERN}"; then
pass=$(( pass + 1 ))
printf '[%s] PRBS %s poll=%d PASS\n' "$(ts)" "${name}" "${n}"
else
fail=$(( fail + 1 ))
printf '[%s] PRBS %s poll=%d FAIL\n' "$(ts)" "${name}" "${n}"
fi
[ "${MAX_POLLS}" -gt 0 ] && [ "${n}" -ge "${MAX_POLLS}" ] && break
[ "${WORKER_STOP}" -eq 0 ] || break
sleep "${POLL_INTERVAL_SEC}"
done
# --- teardown ---
run_step "${unit}" "phydiag ${phy} prbsstat STOp" >/dev/null
run_step "${unit}" "phydiag ${phy} prbs clear" >/dev/null
printf '[%s] PRBS STOP %s polls=%d PASS=%d FAIL=%d\n' \
"$(ts)" "${name}" "${n}" "${pass}" "${fail}"
}
###############################################################################
# Report
###############################################################################
# render_port <name> <unit> <phy> <raw>
render_port() {
local name="$1" unit="$2" phy="$3" raw="$4"
local t polls pass fail verdict
t="$(port_tally "${name}" "${raw}")"
polls="$(printf '%s' "${t}" | awk '{print $1}')"
pass="$(printf '%s' "${t}" | awk '{print $2}')"
fail="$(printf '%s' "${t}" | awk '{print $3}')"
verdict="$(printf '%s' "${t}" | awk '{print $4}')"
if [ "${verdict}" = "SKIP" ]; then
printf -- '----- %s (unit %s, phy %s) : not run this session -----\n\n' \
"${name}" "${unit}" "${phy}"
printf '[%s] RESULT %-12s polls=%-4s PASS=%-4s FAIL=%-4s %s\n\n' \
"$(ts)" "${name}" "${polls}" "${pass}" "${fail}" "${verdict}"
return 0
fi
printf -- '----- %s (unit %s, phy %s) : last %s poll blocks -----\n' \
"${name}" "${unit}" "${phy}" "${TAIL_ENTRIES}"
grep -aE "^-{10} ${name} POLL|PRBS ${name} poll=|prbsstat Ber|^ *[0-9]+ *: " "${raw}" 2>/dev/null \
| tail -n $(( TAIL_ENTRIES * 3 ))
printf '\n'
printf '[%s] RESULT %-12s polls=%-4s PASS=%-4s FAIL=%-4s %s\n\n' \
"$(ts)" "${name}" "${polls}" "${pass}" "${fail}" "${verdict}"
}
write_report() {
local started
started="$(sed -n '1p' "${STATE_FILE}" 2>/dev/null)"
{
printf '#############################################################\n'
printf '[%s] 100G PRBS report\n' "$(ts)"
printf ' started : %s\n' "${started:-unknown}"
printf ' stopped : %s\n' "$(ts)"
printf ' %s : unit %s, phy %s\n' "${PORT_A_NAME}" "${PORT_A_UNIT}" "${PORT_A_PHY}"
printf ' %s : unit %s, phy %s\n' "${PORT_B_NAME}" "${PORT_B_UNIT}" "${PORT_B_PHY}"
printf ' poly=%s prbsstat Interval=%s poll every %ss\n' \
"${PRBS_POLY}" "${PRBS_STAT_INTERVAL}" "${POLL_INTERVAL_SEC}"
printf ' a poll is PASS only when the output contains "%s"\n' "${PASS_PATTERN}"
printf '#############################################################\n\n'
render_port "${PORT_A_NAME}" "${PORT_A_UNIT}" "${PORT_A_PHY}" "${RAW_A}"
render_port "${PORT_B_NAME}" "${PORT_B_UNIT}" "${PORT_B_PHY}" "${RAW_B}"
# Taken after prbsstat STOp + prbs clear, so this is the recovered
# state. During the test the same command would have reported DOWN.
printf -- '--- show interfaces status %s,%s (after prbsstat STOp + prbs clear) ---\n' \
"${PORT_A_NAME}" "${PORT_B_NAME}"
${SUDO} show interfaces status "${PORT_A_NAME},${PORT_B_NAME}" </dev/null 2>&1
} >> "${LOG_FILE}"
}
###############################################################################
# Sub-commands
###############################################################################
prepare_log_on_start() {
local stamp
case "${START_LOG_MODE}" in
append) return 0 ;;
archive)
if [ -s "${LOG_FILE}" ]; then
stamp="$(date '+%Y%m%d-%H%M%S')"
mv -f "${LOG_FILE}" "${LOG_FILE}.${stamp}" || die "cannot archive ${LOG_FILE}"
fi ;;
new)
[ -f "${LOG_FILE}" ] && { : > "${LOG_FILE}" || die "cannot truncate ${LOG_FILE}"; } ;;
*) die "invalid START_LOG_MODE: ${START_LOG_MODE}" ;;
esac
}
# do_start [both|a|b|<port name>]
do_start() {
local which="${1:-${PORTS}}" pids pa="-" pb="-" run_a=0 run_b=0
case "${which}" in
both|BOTH|all) run_a=1; run_b=1 ;;
a|A|"${PORT_A_NAME}") run_a=1 ;;
b|B|"${PORT_B_NAME}") run_b=1 ;;
*) die "unknown port selector '${which}' (use: both | a | b | ${PORT_A_NAME} | ${PORT_B_NAME})" ;;
esac
pids="$(is_running)" && die "already running (pids=${pids})"
command -v "${BCMCMD}" >/dev/null 2>&1 || die "${BCMCMD} not found"
mkdir -p "${LOG_DIR}" || die "cannot create ${LOG_DIR}"
prepare_log_on_start
# Only the selected ports are truncated. A raw with no "PRBS START" line is
# what the report uses to tell "not run this session" from "ran and failed".
[ "${run_a}" -eq 1 ] && : > "${RAW_A}"
[ "${run_b}" -eq 1 ] && : > "${RAW_B}"
printf '%s\n' "$(ts)" > "${STATE_FILE}"
# One worker per selected port; with both, they run at the same time.
# Invoked through bash so the file does not need the execute bit.
if [ "${run_a}" -eq 1 ]; then
setsid bash "${SCRIPT_PATH}" __worker \
"${PORT_A_NAME}" "${PORT_A_UNIT}" "${PORT_A_PHY}" "${RAW_A}" >/dev/null 2>&1 &
pa=$!
disown "${pa}" 2>/dev/null
fi
if [ "${run_b}" -eq 1 ]; then
setsid bash "${SCRIPT_PATH}" __worker \
"${PORT_B_NAME}" "${PORT_B_UNIT}" "${PORT_B_PHY}" "${RAW_B}" >/dev/null 2>&1 &
pb=$!
disown "${pb}" 2>/dev/null
fi
printf '%s\n%s\n' "${pa}" "${pb}" > "${PID_FILE}"
sleep 1
[ "${run_a}" -eq 1 ] && { kill -0 "${pa}" 2>/dev/null || die "${PORT_A_NAME} worker failed, see ${RAW_A}"; }
[ "${run_b}" -eq 1 ] && { kill -0 "${pb}" 2>/dev/null || die "${PORT_B_NAME} worker failed, see ${RAW_B}"; }
printf 'started (%s pid=%s, %s pid=%s)\nlog: %s\n' \
"${PORT_A_NAME}" "${pa}" "${PORT_B_NAME}" "${pb}" "${LOG_FILE}"
# Deliberately NOT printing "show interfaces status" here. With PRBS armed
# the link is out of normal operation and reports DOWN, which reads as a
# failure to anyone glancing at the console. The status is shown in the
# report instead, once PRBS has been cleared.
}
do_stop() {
local pids pa pb i
pids="$(is_running)" || { printf 'not running\n'; rm -f "${PID_FILE}"; return 0; }
pa="${pids% *}"; pb="${pids#* }"
# TERM lets each worker finish its poll and run prbsstat STOp / prbs clear.
[[ "${pa}" =~ ^[0-9]+$ ]] && kill -TERM "${pa}" 2>/dev/null
[[ "${pb}" =~ ^[0-9]+$ ]] && kill -TERM "${pb}" 2>/dev/null
for (( i = 0; i < 120; i++ )); do
kill -0 "${pa}" 2>/dev/null || kill -0 "${pb}" 2>/dev/null || break
sleep 0.5
done
if kill -0 "${pa}" 2>/dev/null || kill -0 "${pb}" 2>/dev/null; then
printf 'workers did not exit in time, sending SIGKILL (PRBS may be left running)\n' >&2
kill -KILL "${pa}" 2>/dev/null
kill -KILL "${pb}" 2>/dev/null
fi
write_report
rm -f "${PID_FILE}"
printf 'stopped\nreport: %s\n' "${LOG_FILE}"
}
do_report() {
local name raw pass fail polls verdict
printf '%-14s %-8s %-8s %-8s %s\n' PORT POLLS PASS FAIL RESULT
printf '%-14s %-8s %-8s %-8s %s\n' -------------- -------- -------- -------- ------
for spec in "${PORT_A_NAME}:${RAW_A}" "${PORT_B_NAME}:${RAW_B}"; do
name="${spec%%:*}"; raw="${spec#*:}"
set -- $(port_tally "${name}" "${raw}")
printf '%-14s %-8s %-8s %-8s %s\n' "${name}" "$1" "$2" "$3" "$4"
done
}
do_status() {
local pids
if pids="$(is_running)"; then
printf 'status : running (%s pid=%s, %s pid=%s)\n' \
"${PORT_A_NAME}" "${pids% *}" "${PORT_B_NAME}" "${pids#* }"
else
printf 'status : stopped\n'
fi
printf 'log : %s\n' "${LOG_FILE}"
do_report
}
usage() {
cat <<EOF
Usage: ${SCRIPT_NAME} {start [both|a|b]|stop|status|report|clear}
start Set up PRBS and poll in the background. With no argument it runs
\$PORTS (currently "${PORTS}"); "a" or "b" runs that port alone,
which is how you tell a genuine port fault from the DUT not coping
with two PRBS streams at once.
start both ports at the same time
start a ${PORT_A_NAME} only
start b ${PORT_B_NAME} only
Port status is NOT shown here: with PRBS armed the link reports
DOWN, which looks like a failure. See the report instead.
stop Stop polling, run prbsstat STOp + prbs clear on both ports, and
render the report into the log
report Print the PASS/FAIL tally per port
status Show worker pids plus the current tally
clear Remove the log and raw captures (must be stopped first)
Log file : ${LOG_FILE}
Raw : ${RAW_A}
${RAW_B}
A poll is PASS only when "phy diag <phy> prbs get" reports "${PASS_PATTERN}".
prbsstat Ber is captured alongside every poll but does not decide the verdict.
A port that was not run in this session reports SKIP, not FAIL.
EOF
}
###############################################################################
# Entry point
###############################################################################
case "${1:-}" in
start) shift; do_start "${1:-${PORTS}}" ;;
stop) do_stop ;;
status) do_status ;;
report) do_report ;;
clear)
is_running >/dev/null && die "still running, stop it first"
rm -f "${LOG_FILE}" "${LOG_FILE}".[0-9]* "${RAW_A}" "${RAW_B}" "${STATE_FILE}"
printf 'log cleared\n' ;;
__worker)
shift
do_worker "$1" "$2" "$3" "$4" ;;
-h|--help|"") usage ;;
*) printf '[ERROR] unknown sub-command: %s\n\n' "$1" >&2; usage; exit 1 ;;
esac
@@ -0,0 +1,203 @@
#!/bin/bash
###############################################################################
# usb_target.sh
#
# Version : V1.1.0
# Author : ETWen
# Date : 20260825
# Purpose : Auto-detect an inserted USB mass-storage device and report either
# its device node, its mount point, or both.
#
# stdout : requested value(s) only, for $(...) capture
# stderr : diagnostic messages
#
# Usage : usb_target.sh [-o dev|mnt|both|id] [-n] [-r seconds]
#
# -o dev print partition device node (e.g. /dev/sda1)
# -o mnt print mount point (e.g. /mnt/usb) [default]
# -o both print "<dev> <mnt>" on one line
# -o id print stable by-id path (/dev/disk/by-id/...)
# -n detect only, do not mount
# -r sec udev enumeration wait, default 15
#
# Exit : 0 success
# 1 no USB mass-storage found
# 2 multiple USB disks detected (refuse to guess)
# 3 no usable filesystem on the device
# 4 mount failed
# 5 mounted but not writable
#
# Version History
# V1.0.0 20260825 Initial Version
# V1.1.0 20260825 Add -o/-n/-r options; expose device node and by-id path
###############################################################################
set -u
MNT_BASE="/mnt/usb"
RETRY=15
OUTPUT="mnt"
DO_MOUNT=1
log() { echo "[usb] $*" >&2; } # diagnostics go to stderr; stdout stays clean
while getopts "o:nr:h" opt; do
case "$opt" in
o) OUTPUT="$OPTARG" ;;
n) DO_MOUNT=0 ;;
r) RETRY="$OPTARG" ;;
h) sed -n '3,30p' "$0" >&2; exit 0 ;;
*) log "invalid option"; exit 1 ;;
esac
done
shift $((OPTIND - 1))
case "$OUTPUT" in
dev|mnt|both|id) ;;
*) log "invalid -o value: $OUTPUT"; exit 1 ;;
esac
# "-o dev" / "-o id" alone does not require mounting.
[ "$OUTPUT" = "dev" ] && [ "$DO_MOUNT" = "1" ] && DO_MOUNT=0
[ "$OUTPUT" = "id" ] && [ "$DO_MOUNT" = "1" ] && DO_MOUNT=0
#------------------------------------------------------------------------------
# Identify the physical disk(s) backing rootfs / /host so they can be excluded.
# Needed because some platforms boot from a USB DOM, which also reports TRAN=usb.
#------------------------------------------------------------------------------
get_system_disk() {
local src
for mp in /host / ; do
src=$(findmnt -no SOURCE "$mp" 2>/dev/null) || continue
lsblk -no PKNAME "$src" 2>/dev/null | head -1
done | sort -u
}
#------------------------------------------------------------------------------
# List candidate USB disks (whole devices, not partitions).
#------------------------------------------------------------------------------
find_usb_disks() {
local sysdisks; sysdisks=$(get_system_disk)
lsblk -dn -o NAME,TYPE,TRAN,RM 2>/dev/null | while read -r name type tran rm; do
[ "$type" = "disk" ] || continue
case "$name" in loop*|ram*|dm-*|sr*|zram*) continue ;; esac
# Older util-linux may not expose the TRAN column; fall back to sysfs.
if [ "$tran" != "usb" ]; then
readlink -f "/sys/block/$name/device" 2>/dev/null | grep -q '/usb[0-9]' || continue
[ "$rm" = "1" ] || continue
fi
echo "$sysdisks" | grep -qx "$name" && { log "skip $name (system disk)"; continue; }
echo "/dev/$name"
done
}
#------------------------------------------------------------------------------
# Pick a mountable partition from a disk; fall back to the whole device for
# superfloppy layouts (filesystem written directly, no partition table).
#------------------------------------------------------------------------------
pick_partition() {
local disk="$1" p
p=$(lsblk -ln -o NAME,TYPE,FSTYPE "$disk" | \
awk '$2=="part" && $3!="" {print "/dev/"$1; exit}')
[ -n "$p" ] && { echo "$p"; return; }
[ -n "$(lsblk -dn -o FSTYPE "$disk")" ] && echo "$disk"
}
#------------------------------------------------------------------------------
# Resolve a device node to a stable /dev/disk/by-id path, if one exists.
#------------------------------------------------------------------------------
resolve_by_id() {
local dev="$1" real link
real=$(readlink -f "$dev")
for link in /dev/disk/by-id/*; do
[ -e "$link" ] || continue
case "$link" in *-part*|*) ;; esac
[ "$(readlink -f "$link")" = "$real" ] && { echo "$link"; return 0; }
done
return 1
}
#------------------------------------------------------------------------------
# Main
#------------------------------------------------------------------------------
disks=""
for i in $(seq 1 "$RETRY"); do
disks=$(find_usb_disks)
[ -n "$disks" ] && break
sleep 1
done
[ -z "$disks" ] && { log "no USB mass-storage found"; exit 1; }
n=$(echo "$disks" | wc -l)
if [ "$n" -gt 1 ]; then
log "multiple USB disks detected, refuse to guess:"
log "$disks"
exit 2
fi
disk="$disks"
part=$(pick_partition "$disk")
[ -z "$part" ] && { log "$disk has no usable filesystem"; exit 3; }
fstype=$(lsblk -no FSTYPE "$part")
log "found $part (disk=$disk fstype=$fstype)"
# --- device-node only: no mount needed ---------------------------------------
if [ "$OUTPUT" = "dev" ]; then
echo "$part"
exit 0
fi
if [ "$OUTPUT" = "id" ]; then
if byid=$(resolve_by_id "$part"); then
echo "$byid"
exit 0
fi
log "no by-id path for $part, falling back to device node"
echo "$part"
exit 0
fi
# --- mount point required ----------------------------------------------------
mp=$(lsblk -no MOUNTPOINT "$part" | head -1)
if [ -z "$mp" ] && [ "$DO_MOUNT" = "0" ]; then
log "$part is not mounted and -n was given"
exit 4
fi
if [ -z "$mp" ]; then
log "mounting $part -> $MNT_BASE"
mkdir -p "$MNT_BASE"
# exfat/ntfs are not always built into the platform kernel.
case "$fstype" in
exfat) grep -qw exfat /proc/filesystems || modprobe exfat 2>/dev/null ;;
ntfs) grep -qw ntfs3 /proc/filesystems || modprobe ntfs3 2>/dev/null ;;
esac
if ! mount -o rw,noatime "$part" "$MNT_BASE" 2>/dev/null; then
log "mount failed (fstype=$fstype)"
exit 4
fi
mp="$MNT_BASE"
# Verify it is actually writable (read-only media, dirty FAT, or full device).
if ! touch "$mp/.wtest" 2>/dev/null; then
log "mounted read-only or no space left"
umount "$mp"
exit 5
fi
rm -f "$mp/.wtest"
else
log "$part already mounted at $mp"
fi
if [ "$OUTPUT" = "both" ]; then
echo "$part $mp"
else
echo "$mp"
fi
exit 0
+11 -1
View File
@@ -5,8 +5,18 @@ project_name = "Blanton"
strTestcase="ENV" ;ENV,EMC,Margin...etc
EN_Margin = 0 ; 1: Enable Margin Test, 0: Disable Margin Test
EN_log = 1 ; 1: Enable log, 0: Disable log
;================================================================
; FAN_SPEED
;================================================================
FAN_SPEED = 100
;================================================================
; Switch Unit Population
;================================================================
SWB_UNIT0 = 1 ; 1 = this DUT has switch unit 0, wait for it ; 0 = skip
SWB_UNIT1 = 1 ; 1 = this DUT has switch unit 1, wait for it ; 0 = skip
;================================================================
; Prompt Definitions
;================================================================
@@ -16,3 +16,8 @@ wait prompt_sonic_root
sendln "show platform leak status"
wait prompt_sonic_root
sendln "show platform leak channels"
wait prompt_sonic_root
sendln "tpm-dut-test fru"
wait prompt_sonic_root
sendln "tpm-dut-test tpm-read"
+8 -5
View File
@@ -1,8 +1,11 @@
; All Event
wait prompt_sonic_root
sendln "date"
sendln "dmesg -T | grep -Ei 'error|fail(ed|ure)?|fault|critical|panic|timeout'"
; i2c event
wait prompt_sonic_root
sendln "dmesg | grep -i error"
sendln "dmesg -T | grep -Ei 'error|fail(ed|ure)?|fault|critical|panic|timeout' | grep -i i2c"
; Clear Event
wait prompt_sonic_root
sendln "dmesg | grep -i fail"
wait prompt_sonic_root
sendln "dmesg | grep -i warning"
sendln "dmesg -C"
+91 -16
View File
@@ -1,30 +1,105 @@
; =============================================================================
; File : utils/wait_init.ttl
; Version : V1.0.0
; Version : V3.0.0
; Date : 2026-08-21
; Author : ETWen
; =============================================================================
; Wait until "show platform temperature" is not "Thermal Not detected".
; Wait until every switch unit ENABLED below reports at least WT_MIN ports up:
; bcmcmd -n <u> -c ps | grep -w up | wc -l
;
; The DUT is not always fully populated. Which units exist is declared ONCE in
; config.ttl, so the traffic blocks in Script A/B/C and this wait agree:
; both units -> SWB_UNIT0 = 1 , SWB_UNIT1 = 1
; unit 0 only -> SWB_UNIT0 = 1 , SWB_UNIT1 = 0
; unit 1 only -> SWB_UNIT0 = 0 , SWB_UNIT1 = 1
; no unit -> SWB_UNIT0 = 0 , SWB_UNIT1 = 0 (bypass, returns immediately)
;
; Enter : prompt of the previous command is NOT consumed.
; Exit : "show platform temperature" is sent, prompt NOT consumed.
; Exit : a command has been sent, prompt NOT consumed (caller does the wait).
;
; Why the "PORTS0=" marker: TTL can only test for a string, so reading a count
; means capturing it. Wrapping the number in an echo gives waitregex a unique
; anchor, and the ECHO of the command cannot false-match -- it reads
; "PORTS0=$(bcmcmd ..." and the pattern requires a digit right after the "=".
;
; Version History:
; V1.0.0 2026-08-21 Initial Version
; V1.0.0 2026-08-21 Initial Version (waited on "show platform temperature"
; until it stopped reporting "Thermal Not detected")
; V2.0.0 2026-08-21 Wait on the bcmcmd port-up count of BOTH switch units
; instead of the thermal sensors; proceed only when both
; are greater than WT_MIN (216).
; V3.0.0 2026-08-21 Add WT_UNIT0 / WT_UNIT1 so a partly populated DUT can
; be declared: wait on the enabled units only, and bypass
; entirely when neither is set.
; V3.1.0 2026-08-21 Take the population from SWB_UNIT0 / SWB_UNIT1 in
; config.ttl instead of local copies, so this wait and the
; traffic blocks cannot drift apart.
; V3.1.1 2026-08-25 WT_INTERVAL 10 -> 60. Fix the WT_MIN comments: the test
; has been `< WT_MIN` (at least) since the threshold was
; corrected, but the text still said "more than", which is
; the off-by-one that once hung Script A.
; =============================================================================
wt_interval = 10
timeout = 30
; ---- which units to wait for: SWB_UNIT0 / SWB_UNIT1, set in config.ttl ------
WT_MIN = 216 ; an enabled unit must report AT LEAST this many ports up
; (216 = 108 loopback pairs x 2, i.e. every cabled port)
WT_INTERVAL = 60 ; seconds between polls
; -----------------------------------------------------------------------------
timeout = 60 ; per-wait cap; bcmcmd ps is not instant
wait prompt_sonic_root
:WAIT_THERMAL
sendln "show platform temperature"
wait "Thermal Not detected" prompt_sonic_root
if result = 2 then
goto THERMAL_OK
; nothing declared -> bypass
if SWB_UNIT0 = 0 then
if SWB_UNIT1 = 0 then
goto NO_UNITS
endif
endif
wait prompt_sonic_root
pause wt_interval
goto WAIT_THERMAL
:THERMAL_OK
:WAIT_PORTS
wt_ok = 1
if SWB_UNIT0 = 1 then
wt_up0 = 0
sendln "echo PORTS0=$(bcmcmd -n 0 -c ps | grep -w up | wc -l)"
waitregex "PORTS0=[ ]*([0-9]+)"
if result = 1 then
str2int wt_up0 groupmatchstr1
endif
wait prompt_sonic_root
if wt_up0 < WT_MIN then
wt_ok = 0
endif
endif
if SWB_UNIT1 = 1 then
wt_up1 = 0
sendln "echo PORTS1=$(bcmcmd -n 1 -c ps | grep -w up | wc -l)"
waitregex "PORTS1=[ ]*([0-9]+)"
if result = 1 then
str2int wt_up1 groupmatchstr1
endif
wait prompt_sonic_root
if wt_up1 < WT_MIN then
wt_ok = 0
endif
endif
if wt_ok = 1 then
goto PORTS_OK
endif
pause WT_INTERVAL
goto WAIT_PORTS
; ---- exits ------------------------------------------------------------------
; Both leave exactly one prompt unconsumed, so the caller's next
; "wait prompt_sonic_root" has something to match.
:NO_UNITS
sendln "echo wait_init:no-switch-unit-declared,bypassing-port-wait"
goto DONE
:PORTS_OK
sendln ""
:DONE
timeout = 0
sendln "show platform temperature"
+63
View File
@@ -0,0 +1,63 @@
echo PORTS0=$(bcmcmd -n 0 -c ps | grep -w up | wc -l)
echo PORTS0=$(bcmcmd -n 1 -c ps | grep -w up | wc -l)
show interfaces status Ethernet513,Ethernet514
//100G Port 513
sudo sfputil lpmode off Ethernet513
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbs set p=3"'
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbsstat STArt Interval=5"'
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbs get"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbsstat Ber"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbsstat STOp"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbs clear"'
//100G Port 514
sudo sfputil lpmode off Ethernet514
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbs set p=3"'
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbsstat STArt Interval=5"'
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbs get"'
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbsstat Ber"'
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbsstat STOp"'
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbs clear"'
//100G Port 513
sudo sfputil lpmode off Ethernet513
sudo sfputil lpmode off Ethernet514
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbs set p=3"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbs set p=3"'
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbsstat STArt Interval=5"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbsstat STArt Interval=5"'
bcmcmd -n 0 -c 'dsh -c "phy diag 268 prbs get"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phy diag 268 prbs get"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbsstat Ber"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbsstat Ber"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbsstat STOp"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbsstat STOp"'
bcmcmd -n 0 -c 'dsh -c "phydiag 268 prbs clear"'
sleep 1
bcmcmd -n 1 -c 'dsh -c "phydiag 268 prbs clear"'
+32
View File
@@ -0,0 +1,32 @@
bcmcmd -n 0 -c '*:port all lb=mac'
bcmcmd -n 0 -c '*:l2 learn off'
bcmcmd -n 0 -c '*:dsh -c "test mode nr=yes"'
bcmcmd -n 0 -c '*:dsh -c "tr 518 Scenario=6 Profile=2 Count=260 PktSize=324 PortList=1-259,270-565 rm=ext CheckLinkTime=20 CheckDataTime=200 testphase=2"'
bcmcmd -n 1 -c "*:port all lb=mac"
bcmcmd -n 1 -c "*:l2 learn off"
bcmcmd -n 1 -c "*:dsh -c 'test mode nr=yes'"
bcmcmd -n 1 -c "*:dsh -c 'tr 518 Scenario=6 Profile=2 Count=260 PktSize=324 PortList=1-259,270-565 rm=ext CheckLinkTime=20 CheckDataTime=200 testphase=2'"
bcmcmd -n 0 -c "*:port all lb=mac"
sleep 1
bcmcmd -n 1 -c "*:port all lb=mac"
sleep 1
bcmcmd -n 0 -c "*:l2 learn off"
sleep 1
bcmcmd -n 1 -c "*:l2 learn off"
sleep 1
bcmcmd -n 0 -c "*:dsh -c 'test mode nr=yes'"
sleep 1
bcmcmd -n 1 -c "*:dsh -c 'test mode nr=yes'"
sleep 1
bcmcmd -n 0 -c "*:dsh -c 'tr 518 Scenario=6 Profile=2 Count=260 PktSize=324 PortList=1-259,270-565 rm=ext CheckLinkTime=20 CheckDataTime=200 testphase=2'"
sleep 1
bcmcmd -n 1 -c "*:dsh -c 'tr 518 Scenario=6 Profile=2 Count=260 PktSize=324 PortList=1-259,270-565 rm=ext CheckLinkTime=20 CheckDataTime=200 testphase=2'"
sleep 1
root@(none):/usr/local/bin# swutil *:l2 learn off
root@(none):/usr/local/bin# swutil *:dsh -c \"test mode nr=yes\"
root@(none):/usr/local/bin# swutil *:dsh -c \"tr 518 Scenario=6 Profile=2 Count=260 PktSize=324 PortList=1-259,270-565 rm=ext CheckLinkTime=20 CheckDataTime=200 testphase=2"
+7
View File
@@ -0,0 +1,7 @@
echo 18-0020 > /sys/bus/i2c/drivers/max31790/unbind
sleep .5
echo 18-0020 > /sys/bus/i2c/drivers/max31790/bind
echo 25-0020 > /sys/bus/i2c/drivers/max31790/unbind
sleep .5
echo 25-0020 > /sys/bus/i2c/drivers/max31790/bind
fan-speed-control.sh 30
+59
View File
@@ -0,0 +1,59 @@
#!/bin/bash
# =============================================================================
# gen_channel_map.sh - Build docs/LTC2980_channel_map.csv from settings/*.conf
# Version : V1.0.0
# Date : 2026-08-27
# Author : ETWen
# =============================================================================
# Runs on the DEV machine, not the DUT. Reads every LTC2980 margin settings
# file and flattens CH<n>_VNOM / CH<n>_NET into one CSV:
#
# Board,CONN,Ch,NetName,Vnom
# CB,CONN13,CH0,V5P0_ALW,5.0
#
# NC channels keep their literal NC / "-" so the page numbering still lines up
# (ch/8 -> CHIPS[] index, ch%8 -> LTC2977 PAGE). Do not filter them out.
#
# Usage:
# tools/gen_channel_map.sh # write docs/LTC2980_channel_map.csv
# tools/gen_channel_map.sh - # write to stdout
# =============================================================================
set -eu
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
SET_DIR="${REPO_ROOT}/src/Script_ABC_Blanton/Blanton_Script/LTC2980_Margin_Script/settings"
OUT="${1:-${REPO_ROOT}/docs/LTC2980_channel_map.csv}"
[ -d "${SET_DIR}" ] || { echo "settings dir not found: ${SET_DIR}" >&2; exit 1; }
# 3-argument match(str, regex, array) is a gawk extension. Under mawk it is a
# syntax error at best and an empty CSV at worst - fail loudly instead.
awk 'BEGIN { if (match("a", /a/, m) != 1) exit 1 }' 2>/dev/null \
|| { echo "need gawk (3-arg match); this awk is: $(awk --version 2>&1 | head -1)" >&2; exit 1; }
gen() {
printf 'Board,CONN,Ch,NetName,Vnom\n'
# CB first, then SWB0/SWB1 CONN13..CONN16 - matches the physical walk order.
for f in "${SET_DIR}"/Blanton_CB_CONN*.conf \
"${SET_DIR}"/Blanton_SWB0_CONN*.conf \
"${SET_DIR}"/Blanton_SWB1_CONN*.conf ; do
[ -f "${f}" ] || continue
base="$(basename "${f}" .conf)" # Blanton_SWB0_CONN13
board="$(echo "${base}" | cut -d_ -f2)" # SWB0
conn="$(echo "${base}" | cut -d_ -f3)" # CONN13
# .conf files are LF per .gitattributes, but strip CR anyway - a stray
# CR would land inside the net name and travel into the CSV unseen.
sed 's/\r$//' "${f}" | awk -v b="${board}" -v c="${conn}" '
match($0, /^CH([0-9]+)_VNOM="([^"]*)"[^"]*CH[0-9]+_NET="([^"]*)"/, m) {
printf "%s,%s,CH%s,%s,%s\n", b, c, m[1], m[3], m[2]
}'
done
}
if [ "${OUT}" = "-" ]; then
gen
else
gen > "${OUT}"
echo "wrote ${OUT} ($(($(wc -l < "${OUT}") - 1)) channels)"
fi