On September 10, I was shooting App Store screenshots for Amana. The collection screen shows every sky you've photographed, each with the name the app gave it. I fed seven photos through the English UI: a mountain ridge, a blue sky with cumulus, a pink townscape, orange power lines, a rice field, navy and gold, the sea. All seven went in from the photo library at half past four, so the engine treated every one as an autumn evening. Seven skies, three names between them:
Day's End
Word from the West
Day's End
Day's End
Word from the West
Word from the West
The Quick Autumn Dusk
The screenshot script already had a duplicate check, because a blue sky and a sunset over mountains had come out with the same name the day before. Its advice was "swap a photo and reshoot." I swapped photos three times and got a different pair each time. The English collection shot ended up with three cards and an empty lower half.
The unit suite was 192 tests in 26 suites when I ran it, all green, and nothing in it changed when the fix went in. This post is about why green told me nothing.
Naming is a pure function: offline, no model, no randomness. The view model averages the image down to one color with CIAreaAverage, reads hue, saturation and brightness, and derives a cloudiness estimate (low saturation plus high brightness reads as "white and overcast"). Then:
named = NamingEngine.name(hue: features.hue,
hour: hour, // from the photo's date
month: month,
condition: condition,
cloudiness: features.cloudiness,
language: language)
Four numbers. GPS never reaches this function. Neither, in practice, does the weather: the condition parameter exists, but the capture screen never passes it, so it's always .unknown. Two photos of the same sky from different cities are the same input.
Inside, each continuous value is quantized into a band: four time bands (dawn, day, dusk, night), three cloud bands, five hue bands named after Japanese colors. A nested switch over the bands returns a cell, and each cell holds a pool of two or three hand-written names. Two photos that land in the same bands hit the same cell.
That was already a problem in June. Similar skies kept getting the same name, so I added a spread index computed from the raw values before quantization:
static func spreadIndex(hue: Double, hour: Int, cloudiness: Double) -> Int {
let h = normalizedHue(hue) // 0..<360
let c = min(max(cloudiness, 0), 1) // 0...1
let hr = Double(((hour % 24) + 24) % 24) // 0...23
// Mix the features at different scales so nearby inputs land on different candidates.
let raw = (h + c * 100 + hr * 7).rounded()
return Int(raw)
}
// the pick: pool[((spread % n) + n) % n]
That commit doubled the vocabulary from 46 to 88 names and added a test, variantSpreadsWithinSameBand. It walks hue in 1.5° steps through one band and asserts that the resulting set of names has count >= 2. It passed in June. It was still passing in September.
In July a seasonal layer went on top: a word bank of entries tagged by season, time of day (five bands) and sky pattern, each with a Japanese and an English name. The lookup filters the bank for the photo's context and picks with the same spread index:
case .japanese:
return pickEntry([base] + pool.map(\.ja), spread)
case .english:
return pool.isEmpty ? base : pickEntry(pool.map(\.en), spread)
Note the asymmetry. Japanese gets the original engine's answer plus the bank. English gets only the bank, because the 88 originals have no translations. For "autumn × evening glow," three bank entries matched: The Quick Autumn Dusk, Day's End, Word from the West. Seven photos, three names. The pigeonhole principle guarantees at least four collisions before any arithmetic runs.
Still, shouldn't the spread index at least distribute seven photos across three names? Look at what actually varies. Hue and cloudiness are per photo. The hour is not: all seven went through the picker within the same hour, so hr * 7 was 112 for every one of them. (Amana takes the hour from EXIF DateTimeOriginal for library photos and from the clock for camera shots. My staging script had stripped the EXIF from these seven to control their order in the picker, so they fell back to the clock. Either way: one shoot, one hour.)
Because 7·hour is an integer, rounding commutes with it, and the index is exactly r + 112 where r = round(hue + 100·cloudiness). The pick is (r + 112) mod 3. Adding a constant modulo n is a bijection on the residues: it rotates the names, it never merges or splits them. The number of distinct names is decided entirely by {r mod n}. Here are the seven values from the harness described below, and their residues:
r = 81, 274, 264, 96, 370, 289, 251
r mod 3 = 0, 1, 0, 0, 1, 1, 2
Three classes. Change the hour to 17 or 18 and every photo moves together: the same three names, reassigned. That's why reshooting later didn't help, and why swapping one photo just produced a different pair. For two photos from different hours the shift does differ, so the term isn't useless. It just can't separate photos that share an hour, which is exactly what one evening outside produces.
The intent in that comment, "mix the features at different scales," was reasonable. Hour was being mixed in at scale 7. Scale is irrelevant for a value that's constant across the batch.
The fix has to be more words. How many? The tempting answer is "add some and reshoot," but the shoot window is 16:00–18:30 once a day, and I'd already burned three attempts.
Instead I built a small harness: the feature-extraction code copied verbatim, linked against the real NamingEngine sources, compiled with swiftc for the iOS simulator and run with simctl spawn against the seven JPEGs. A Python approximation wouldn't do, because CIAreaAverage with a null working color space doesn't match PIL's averaging. First I ran the harness with the old vocabulary. It reproduced the three English names and the five Japanese names from the shoot exactly, and only at hour 16, which matches the shoot log. That's the positive control; without it the numbers below would be a guess.
With r in hand, the rest is arithmetic. Count the distinct residues for each candidate pool size:
|
Words in the pool |
Distinct names for the 7 photos |
|---|---|
|
3 (shipped) |
3 |
|
8 |
4 |
|
9 |
6 |
|
10 |
5 |
|
11 |
7 |
|
12 |
5 |
It's not monotonic. Eight words, the pool after my first pass had added five, gave four names. Ten is worse than nine. Eleven was the first size where all seven separated, so the second pass added three more entries, scoped to autumn only so the other seasons' pools stayed where they were. All eight additions are appended at the end of the bank; existing entries keep their positions, and the diff is 96 lines added, none removed.
Japanese went from five names to five, which was expected. Its pool starts with the original engine's answer at index 0, and those were seven different names for these seven photos; the collisions were among the bank words. It also has a trap the English side doesn't: if a new Japanese bank word happens to equal an original engine name, two photos silently merge into one card and nothing fails. I checked the new words against the originals before adding them.
@Test func autumnEveningGlowEnglishPoolIsWideEnough() {
for sky in NamingEngine.SkyPattern.allCases {
let names = Set(NamingWordbank.matches(season: .autumn, time: .eveningGlow, sky: sky)
.map(\.en.name))
#expect(names.count >= 7) // message trimmed
}
}
It's a floor on the pool, not a promise of seven distinct names. Seven other photos with other residues can still collide at eleven words, and the sweep above shows twelve words would have given these seven photos only five names. What the test does is fail loudly if the pool ever shrinks, and it names the actual cell that broke, which none of the earlier tests did.
Before this, the naming engine had four kinds of test: determinism across a grid (same input, same output), no forbidden characters, non-empty output, and diversity, meaning the whole grid produces at least 30 distinct names and a walk within one band produces at least two. All true. All still true. None of them asks the question a user asks: given the handful of photos I took this evening, how many different names will I see?
"At least two" was chosen when the pools had two entries, so it was the strongest assertion available. It then quietly became the weakest one as the vocabulary grew around it. And a determinism test is in some sense the opposite of a diversity test: it proves the function is a function, not that its image is large.
What I'd tell my June self: whenever a hash or spread function is supposed to create variety, write down what varies per item and what is constant across the batch, then test the distribution over a realistic batch rather than the existence of two values. The math here was one line. Seeing that I needed the math took three reshoots.
Amana is on the App Store. The names are still deterministic; there are just more of them now.