summaryrefslogtreecommitdiff
path: root/stacks/tncs/docs/TNC-RECOVERY.md
blob: 2e36357a8c56f05a44a26f6e416c09a50597f62d (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
# TNC recovery (software-first)

TheFirmware / TNC-2 class devices (Landolt TNC2C, PK-TNC2, TAPR TNC-2) can usually return to **host command mode** and **KISS** without a power cycle. Power-cycle with DTR held high is a **rescue fallback** only.

Shared implementation: `stacks/tncs/tnc_serial_recovery.py` (used by `tnc2c-host-reset`, `tnc2c-boot-wait`, and `max25d` `kiss_bridge`).

## When to use what

| Situation | First action | Rescue (if software fails) |
|-----------|--------------|----------------------------|
| Echo-only (`INFO` → `INFO`, no banner) | `max25d` serial watch / `./tnc2c-host-reset.sh` | `./tnc2c-boot-wait.sh` only if cold-boot without DTR |
| After `max25d` prep `error-host` | Auto boot-wait escalate (power OFF 10s → ON) | `./tnc2c-boot-wait.sh` if escalate disabled |
| PK-TNC2 (9600 8N1) | `./pktnc2-boot-wait.sh --recover-only` | `./pktnc2-boot-wait.sh` + power cycle |
| Cold boot, never saw `cmd:` | `./tnc2c-boot-wait.sh` (DTR before power-on) | Fix wiring (DTR/RTS, CTS bridge) |

## Software recovery ladder (TheFirmware TF 2.7 native)

Run with **DTR+RTS high** and the port open:

1. Passive listen (~1.5 s) + **DTR settle 2 s** after open (max25d)
2. KISS return `0xC0 0xFF 0xC0` (firmware reset → banner)
3. Buffer flush `^Q^X` + JHOST 0 (leave host mode)
4. **ESC V** — version / terminal probe
5. **ESC QRES** — software cold boot (no mains power cycle; DTR must stay high)
6. ESC E 0 + second ESC QRES
7. Legacy TAPR `kiss off` + `INFO` (last resort only)

Success markers: `TheFirmware`, `NORD`, `Version 2.7`, `Checksum`, `cmd:`.

Then enter KISS: **ESC `@K`** (`0x1B 40 4B`). MYCALL: **ESC I** `<call>`.

Native TF sequence: see [Software recovery ladder](#software-recovery-ladder-thefirmware-tf-27-native) above.

## Operator commands

```bash
cd stacks/tncs

# Software recovery (no power cycle)
./tnc2c-host-reset.sh              # terminal mode
./tnc2c-host-reset.sh --kiss       # terminal + KISS
./tnc2c-boot-wait.sh --recover-only

# PK-TNC2
./pktnc2-boot-wait.sh --recover-only

# Rescue: power cycle with DTR held (Landolt TNC2C cold-boot requirement)
./tnc2c-boot-wait.sh               # power OFF → ON while script runs
```

## MAX25 + HyBBX contract

| Layer | Responsibility |
|-------|----------------|
| `max25d` / boot-wait | DTR sequencing, software recovery, `kiss on` / `auto` |
| HyBBX `packet_radio` | Attach only — `kiss_entry = none` |
| Power cycle | Rescue when DTR was low at cold boot or hardware hang |

HyBBX INI (production): `kiss_entry = none`, `persist = 255` (CB CSMA), optional `[max25] check = yes`.

max25d runs `recover_terminal()` in `kiss_bridge.stabilize_session()` before `MYCALL` and KISS entry. The KISS RX thread starts **only after** successful stabilization (avoids a race where RX consumes recovery replies).

### max25d serial watch (automatic)

When `[stack] serial_watch = yes` and `stack_recover_only = yes` (default), max25d **opens the TNC serial port itself** — no `boot-wait` subprocess (avoids port conflict). Recovery runs inline via `stabilize_session`.

| INI key | Default | Role |
|---------|---------|------|
| `serial_watch` | yes | Enable periodic probe + auto-repair |
| `serial_watch_interval` | 60 | Seconds between health probes |
| `serial_repair_cooldown` | 20 | Minimum gap between repair attempts |
| `stack_recover_only` | yes | Daemon start uses `--recover-only` (no power cycle) |
| `stack_retry_interval` | 120 | Retry failed stack prep with recover-only |
| `serial_bootwait_escalate` | yes | Escalate to boot-wait after repeated inline failures |
| `serial_bootwait_escalate_after` | 3 | Inline `error-host` failures before boot-wait |
| `serial_bootwait_escalate_cooldown` | 300 | Minimum seconds between boot-wait escalations |

Triggers: initial prep `error-host` (immediate boot-wait escalate), `error-host` / `error-kiss` / `error-tx` / `error-io`, failed TX (one auto-retry). Serial watch still escalates after `serial_bootwait_escalate_after` consecutive inline failures if prep did not already escalate.

**Healthy KISS:** when the backend is `ready`, serial watch does **not** tear down KISS on the periodic interval — repair runs only when status is in the error set above (not a periodic `kiss off` / probe on a working session).

## Firmware RX diagnostics (max25d / boot-wait)

On each recovery step, logs include byte count, matched firmware markers (`TheFirmware`, `cmd:`, …), and a printable preview. On failure:

- `recovery: firmware assessment — …` — classified state (silent, echo-only, binary/KISS, non-banner text)
- `recovery: RX capture — …` — full accumulated RX summary with hex snippet

Use these lines to distinguish **DTR/cold-boot** (echo-only, 0 B passive) from **wrong baud/line** (garbage/binary) from **stuck KISS** (binary frames, no `cmd:`).

Use when:

- Software ladder ends in echo-only or silence
- TNC was powered on **without** DTR high (Landolt TNC2C terminal detection)
- Suspected hardware hang

```bash
./tnc2c-boot-wait.sh
# While script listens: power OFF 10 s, power ON — keep script running (DTR stays high)
```

Do **not** close the serial port between boot-wait and HyBBX/max25d attach — DTR drop can return the TNC to echo mode.

## See also

- [TNC2C-OPERATIONS.md](TNC2C-OPERATIONS.md) — daily ops
- [HYBBX-TNC2C.md](HYBBX-TNC2C.md) — HyBBX attach
- [../../docs/PACKET-RADIO.md](../../docs/PACKET-RADIO.md) — serial profiles, CSMA
git clone -b <branch> https://cgit.mode42.com/<repo>.git
git clone -b <branch> git://cgit.mode42.com/<repo>.git

info@mode42.com