# Chapter 6 - Passwords, SHA-256, and Keyed Security Design (Weeks 6-7)

<!-- lang-switch -->
> [🌐 中文版](https://yukinoshita-lin.github.io/nsf5-steganography/zh/content/ch06.html)




> **Try it |** Goals: understand why embedding positions must depend on a key; explain decoder self-synchronization and tamper perception; know the limits of this design. Two real-computed figures (Fig. 6-1 / Fig. 6-2) are included.

## 6.1 What If the Embedding Positions Are Fixed?

Suppose you always embed sequentially from the top-left. Knowing the algorithm, an attacker can: ① read the LSBs of the first N pixels and recover the message; ② copy the message in place onto another "lookalike" image (position-replacement attack); ③ perturb a fixed region to destroy your message while leaving the rest untouched.

So modern steganography makes the embedding path a **key-controlled pseudo-random order**: without the password (or the content hash), nobody knows where the message is, and positional attacks are infeasible.

![Fig. 6-1 keyed permutation path](../assets/hash_permute_path.png)

*Fig. 6-1 (real computation: using the project's `permute_index()` splitmix64 + Fisher-Yates, drawn as "visit order" on a 64x64 grid. Left is password A, right is password B - the two "orders" are entirely different, both looking like random noise)*

> **Reading the figure |** Brighter color = visited earlier. Change the password and the entire order changes - an attacker without the password cannot predict "where the message hides". That is the meaning of **keying**: **security comes from "not knowing the positions", not from "being invisible".**

## 6.2 Two Seed Rules: Recognize the Image First, Then Find the Payload

The project splits pixels into two pools: the **header pool** (the first N_h blocks) and the **body pool** (the rest). Key derivation uses two independent seeds:

- **seed0 = derived from the password only**: determines the header-pool shuffle order. The header embeds cover_hash - the first 16 bytes (128 bits) of a SHA-256 of the original image bytes;
- **seedB = cover_hash + password combined**: determines the body-pool shuffle order and block partition.

To decode, the receiver first rebuilds seed0 with their own password, reads back cover_hash from the header, then rebuilds seedB together with the password to locate the payload. A wrong password or a damaged header instantly breaks the subsequent path - that is what `test_pwd_mismatch()` checks.

*Seed derivation in the project (real code)*

```python
def derive_seed(image_bytes, password=""):
    base = hashlib.sha256(image_bytes).hexdigest()
    if password:
        base = hashlib.sha256(
            base.encode() + password.encode()).hexdigest()
    return int(base, 16)
```

## 6.3 Deterministic Shuffle: splitmix64 + Fisher-Yates

A seed alone is not enough; it must become "a deterministic permutation of 1..N". The project uses splitmix64 as the pseudo-random source and runs Fisher-Yates: from the end forward, each step swaps with a position chosen by a pseudo-random number. The same seed always yields the same permutation - that is the mathematical basis of "self-synchronization".

![Fig. 6-2 keyed permutation scatter](../assets/hash_permute_scatter.png)

*Fig. 6-2 (real computation: x-axis = pixel position, y-axis = visit rank. Password A (blue) and password B (red) are each a permutation that uses every point once; the grey dashed line is the no-shuffle order (directly predictable = dangerous). The same password is always the same curve -> self-sync; a different password scatters it entirely)*

> **Back to the code |** Open `permute_index()` in `src/ns5_core.py`: when the DLL is available it calls C++ `nsf5_permute`, otherwise it uses the same-algorithm Python fallback. **Both implementations must give exactly the same permutation**, otherwise an image embedded with the DLL cannot be decoded on a DLL-less machine. `src/cppembed.py`'s `selfcheck()` does exactly this cross-language consistency check.

## 6.4 Tamper Perception: What It Can Actually Sense

The embedder reports `cover_hash` (the original image hash) in its report. If any pixel changes, the image bytes' SHA-256 changes. Recomputing the hash of the received image and comparing it with the reported (or header-decoded) hash: mismatch = the image was altered. Better still, since the body path depends on this hash, even a one-pixel change shifts the body shuffle seed, and decoding immediately breaks or turns into garbage.

Be honest about the limits: this is **keying and tamper perception**, not cryptographic **integrity authentication**. The header hash is not MAC'd with a key; in theory an attacker able to re-embed the whole image can also update the header hash. The paper and README list this as future work (e.g. a keyed MAC), so distinguish "detecting ordinary tampering" from "resisting malicious re-embedding".

> **Tip |** One sentence distinction: **tamper perception** = "I see it was altered"; **integrity authentication** = "I can prove it was not altered and an attacker cannot cover it up". The former detects, the latter resists - different strengths.

## 6.5 Hands-On: Wrong Password and Pixel Tampering

*Experiment A: a mismatched password should fail; Experiment B: decode after changing 1 pixel*

```python
import sys; sys.path.insert(0, "src")
import numpy as np; import image_io as IO
from ns5_core import embed_string, extract_string, get_image_hash
 
img = IO.load_as_gray("img/cover.png")
stego, rep, _ = embed_string(img, "secret", method="nsF5",
                             p=3, password="A")
try:                       # wrong password: the bytes decode to garbage and the length check fails
    print(extract_string(stego, method="nsF5", p=3, password="B"))
except ValueError as e:
    print("wrong password -> decode failed:", e)
tampered = stego.copy(); tampered[0, 0] ^= 1   # change only 1 pixel
print(get_image_hash(stego)[:16], "vs", get_image_hash(tampered)[:16])
try:                       # after tampering the hash no longer matches and decoding fails
    print(extract_string(tampered, method="nsF5", p=3, password="A"))
except ValueError as e:
    print("image tampered -> decode failed:", e)
```

> **Watch out |** Experiment B may fail as a ValueError, garbage, or a truncated message, depending on which block the changed pixel falls in. Do not expect "always an exception" - tamper perception's engineering behavior is: the hash comparison fails and the payload cannot self-synchronize.

## 6.6 Summary and Self-Check

- Fixed path = predictable = can be copied and targeted; a keyed path makes the hiding positions depend on the password + content (Fig. 6-1);
- The header holds a 128-bit content hash, and the body seed depends on it -> decoding needs no external sync;
- A deterministic permutation (splitmix64 + Fisher-Yates) is the mathematical guarantee of self-sync (Fig. 6-2);
- Keying is not integrity authentication; strong adversarial scenarios need a keyed MAC.

> **Think about it |** If two different original images are used to embed the same message, will the hidden positions in the two stego images be the same? Why? (Hint: what determines the body seed?) One layer deeper: since security relies on "not knowing the position", could an attacker "guess" the shuffle rule via statistics? (See Chapter 8 - that is exactly detection's entry point.)
