feat(prbs): Add 100G PRBS monitor; fix the job log path and USB mount
port_prbs_monitor.sh drives PRBS on both 100G uplinks: lpmode off, arm the pattern, then poll `phy diag <phy> prbs get` in the background until stopped, counting a poll as PASS only when the output says "PRBS OK!". prbsstat Ber is captured beside every poll for the record but does not decide the verdict. `stop` runs prbsstat STOp and prbs clear before rendering the report, so the test does not leave PRBS armed. `start a` / `start b` runs one uplink alone, which is how a genuine port fault is separated from the DUT not coping with two PRBS streams at once. A port that was not run reports SKIP rather than FAIL -- deciding on poll count alone would have made a deliberate single-port run look like half the hardware was broken. The port status is deliberately not printed at `start`: with PRBS armed the link reports DOWN, which reads as a failure to anyone glancing at the console. It appears in the report instead, after prbs clear and labelled as the recovered state. The calls in Script A, B and C are committed but commented out -- PRBS is still under bring-up. Fix: the job directory was written as /host/hw-eval/current/jobs in some places and /host/hw-eval/jobs in others, so what Script B wrote was not what Script C collected. Unified on /host/hw-eval/jobs. Fix: Script C copied to /mnt/usb assuming it was mounted. It now finds the device with usb_target.sh and mounts it first, so the logs actually leave the DUT. Script C also reorders its results -- stress logs, then the MGMT ping report, then 100G status -- and cats the bgctl ssd/usb logs. Script A -> V1.0.9. Docs and docs/release-notes/v1.0.9.md updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RetWKFZFG1ZHcitQAyhwM
This commit is contained in:
+17
-7
@@ -40,8 +40,9 @@ LTC2980 margin 控制器。TTL 只負責「按順序 `sendln` 並把畫面收進
|
||||
| BMC 存取 | host 端 `bmc-manager run "<cmd>"`(不需對 BMC 開 SSH,繞開其 busybox 工具鏈);`bmc-first-enroll` / `bmc-manager version\|status` |
|
||||
| 背景工作管理 | `bgctl run/list/stop --all/reset`(平台工具);長時間 ping 用 `systemd-run --unit=` 掛成 transient unit |
|
||||
| 管理網路 | iproute2(`ip link set` / `ip address replace` / `ip route replace ... metric`)+ `ping -I` |
|
||||
| 測試 session | `hw-test-session start\|log\|status\|finish`;job log 收在 `/host/hw-eval/current/jobs/` |
|
||||
| 測試 session | `hw-test-session start\|log\|status\|finish`;job log 收在 `/host/hw-eval/jobs/` |
|
||||
| 風扇控制 | max31790 driver rebind(`/sys/bus/i2c/drivers/max31790/{bind,unbind}`)+ `fan-speed-control.sh <%>` |
|
||||
| 100G PRBS | `bcmcmd -n <unit> -c 'dsh -c "phy diag <phy> prbs …"'`(set / prbsstat STArt / get / Ber / STOp / clear) |
|
||||
| 壓力測試 | `mlucas-avx2`(AMD CPU)、`bgctl run` 排程的 `memtester` / `qfx5252-stress-ssd` / `qfx5252-stress-usb` |
|
||||
| 資料面流量 | `bcmcmd`(Broadcom drivshell)VLAN loopback + `tx` burst,switch unit 0 / 1 |
|
||||
| 流量報表 | DUT 上 awk(`blanton_traffic_linespeed report`);離線 `python3 tools/bcm_mibpair_report_V1.1.0.py` |
|
||||
@@ -177,7 +178,10 @@ Blanton_TTL_Script/
|
||||
├── blanton_traffic_linespeed.sh # SWB loopback 線速流量(bcmcmd unit 0/1)
|
||||
│ # init/clear/show/start/stop/ps/report/run
|
||||
├── mgmt_ping_monitor.sh # 10G/1G 管理網路 ping 監控(iproute2)
|
||||
│ # eth0→TARGET_10G、eth1→TARGET_1G,兩張卡都保持 up
|
||||
│ # eth0→TARGET_10G、eth1→TARGET_1G,兩張卡同時持續 ping
|
||||
├── port_prbs_monitor.sh # 100G uplink PRBS 測試(bcmcmd phy diag)
|
||||
│ # ⚠️ 已納入交付,但 A/B/C 的呼叫仍註解掉(bring-up 中)
|
||||
├── usb_target.sh # 偵測 USB 裝置節點/掛載點(排除系統碟)
|
||||
├── bmc_monitor.sh # BMC 監控:free -m + I2C 寫入/讀回圖樣測試
|
||||
├── bmc_monitor_ddr.sh # ⚠️ 已停用:B/C 內呼叫處已註解,DDR 壓力改走 bgctl + bmc-manager
|
||||
│ # ↑ 三支都要 chmod +x 才能跑(見 Key Constraints)
|
||||
@@ -314,6 +318,8 @@ PROFILE_CHANNELS=( "0:high:8" "1:low:8" ... ) # "channel:operation:change_perc
|
||||
| **管理網路 ping** | `mgmt_ping_monitor.sh {start\|stop\|status\|fg\|tail\|clear\|summary}` | 一輪 = 10G leg + 1G leg,各自 `ip address/route replace` 後 `ping -I <if>` 打自己的對端;兩張卡都保持 up。每 leg 記 `ip -s link show` 與該卡的 TX/RX packet delta,另出一行可 grep 的 `RESULT` |
|
||||
| **BMC 監控** | `bmc_monitor.sh {start\|stop\|status\|fg\|tail\|clear}` | 由 host 端 `bmc-manager run` 週期取樣 `free -m`,另加一組 **I2C 寫入/讀回圖樣測試**(bus 1、addr 0x41、offset 0x00C0,寫 `55AA55AA` 讀回、再寫 `AA55AA55` 讀回)。B 啟動、C 停止並 `cat log/bmc_poll.log` |
|
||||
| **BMC DDR 壓力** | `bmc_monitor_ddr.sh` | ⚠️ **已停用** —— B/C 內呼叫處已註解,改由 `bgctl run bmc-manager run memtester` 取代。檔案保留未刪 |
|
||||
| **100G PRBS** | `port_prbs_monitor.sh {start [both\|a\|b]\|stop\|status\|report\|clear}` | 兩個 uplink 各一個 worker,可同時或單獨跑。setup(`lpmode off` → `prbs set` → `prbsstat STArt`)→ 週期 `prbs get`(判定)+ `prbsstat Ber`(只記錄)→ teardown(`STOp` + `prbs clear`)。輸出含 `PRBS OK!` 才算 PASS。**⚠️ A/B/C 的呼叫目前註解掉** |
|
||||
| **USB 目標偵測** | `usb_target.sh [-o dev\|mnt\|both\|id] [-n] [-r sec]` | 找出插入的 USB 儲存裝置,排除 `/` 與 `/host` 的底層碟(本平台可能從 USB DOM 開機),多顆時拒絕猜;掛載時會驗證真的可寫 |
|
||||
| **打包交付** | `publish/publish.sh [-v Ver] [-f] [-z] [-n]` | 版號讀 Script A 檔頭,複製成 `publish/Script_ABC_Blanton_<Ver>/`,排除 `*.log`,可另出 zip |
|
||||
|
||||
---
|
||||
@@ -330,7 +336,7 @@ Script A ── logopen Logs/Blanton_Margin_<ts>.log ──┐ (之後所有畫
|
||||
│ │
|
||||
├─ 登入 admin → sudo -i → root@sonic:~# │
|
||||
├─ date -s "<host 時間>" ← 讓 DUT 時戳可對照
|
||||
├─ rm /host/hw-eval/current/jobs/* ← 清掉上一輪的 job log
|
||||
├─ rm /host/hw-eval/jobs/* ← 清掉上一輪的 job log
|
||||
├─ hw-test-session start / log / status
|
||||
├─ 風扇:max31790 unbind→sleep→bind(18-0020 與 25-0020 各一次)
|
||||
│ → fan-speed-control.sh <FAN_SPEED> → 回答 y 確認 │
|
||||
@@ -374,7 +380,8 @@ Script C
|
||||
├─ nvme smart-log / smartctl -x /dev/nvme0
|
||||
├─ traffic: stop → report ← per-pair TX/RX 交叉比對,PASS/FAIL/NA(exit 2 = 有 FAIL/NA)
|
||||
├─ show reboot-cause / show uptime ← soak 期間有沒有意外重開,一眼看得到
|
||||
├─ cp -r /host/hw-eval/current/jobs/ /mnt/usb/jobs-<YYYYmmdd>_<HHMM>
|
||||
├─ usb_target.sh -o dev → sync → mount <dev> /mnt/usb
|
||||
├─ cp bmc_poll.log / mgmt_ping.log → jobs/,再 cp -r jobs/ /mnt/usb/jobs-<YYYYmmdd>_<HHMM>
|
||||
└─ hw-test-session finish
|
||||
▼
|
||||
[messagebox "GOOD JOB! Test Case DONE"] → 人工檢查 log
|
||||
@@ -717,14 +724,17 @@ margin status、全部 VRM 軌的 Vin/Vout/Iout/Temp、CB + ICB 溫感讀值,
|
||||
- **`bmc_monitor_ddr.sh` 已停用**:B/C 內的呼叫已註解,DDR 壓力改由 `bgctl run bmc-manager run memtester`
|
||||
取代。檔案還留在 repo,待確認不再需要後刪除,或在檔頭標註停用。
|
||||
(`bmc_monitor.sh` 已於 2026-08-25 重新啟用,並加上 I2C 寫入/讀回圖樣測試。)
|
||||
- **`rm /host/hw-eval/current/jobs/*` 清不掉子目錄**:Script A 用的是不帶 `-r` 的 `rm`,但 Script C
|
||||
- **`rm /host/hw-eval/jobs/*` 清不掉子目錄**:Script A 用的是不帶 `-r` 的 `rm`,但 Script C
|
||||
是 `cp -r` 那個目錄,代表裡面可能有子目錄。清不乾淨的話,上一輪的 job log 會被當成這輪的一起
|
||||
複製到 USB。要確認 `bgctl` 的產出結構,必要時改成 `rm -rf .../jobs/*`。
|
||||
- **100G PRBS 尚未接進 A/B/C**:`port_prbs_monitor.sh` 已隨包交付、可手動執行,但三支腳本裡的呼叫
|
||||
都還註解著(bring-up 中)。要啟用時記得 PRBS 期間 link 會顯示 down,`stop` 之後才會回來。
|
||||
- **Script C 的 USB 掛載走 `mount` 而非 `usb_target.sh -o mnt`**:目前是 `-o dev` 取節點再自己
|
||||
`mount <dev> /mnt/usb`,繞過了工具內建的 `mkdir -p`、exfat/ntfs `modprobe` 與可寫性驗證。
|
||||
掛載點不存在、檔案系統模組沒載、或媒體唯讀時只會失敗一行,job log 就沒帶出來。
|
||||
- **風扇確認提示是無限等待**:`wait "Set all configured fan channels to"` 之後才 `sendln "y"`;
|
||||
若該版 `fan-speed-control.sh` 不再詢問(或字串改了),這個 `wait` 會永遠等下去,Script A 停在那裡
|
||||
且不報錯。加個 `timeout` + 分支會比較保險(記得還原 `timeout = 0`)。
|
||||
- **USB 掛載點未檢查**:Script C 直接 `cp -r ... /mnt/usb/`,`/mnt/usb` 沒掛載時只會失敗一行,
|
||||
整輪測試的 job log 就沒被帶出來。可在複製前先 `mountpoint -q /mnt/usb` 判斷。
|
||||
- **Script C 沒有倒出 `mgmt_ping.log`**:C 會 `stop` 管理網路監控並印 `ip -s link show`,
|
||||
但沒有 `cat ~/Blanton_Script/log/mgmt_ping.log`,所以整個 soak 期間累積的 per-leg `RESULT`
|
||||
留在 DUT 上、沒進 master log。Script A 那輪基線有 `cat`。
|
||||
|
||||
Reference in New Issue
Block a user