Introduction
Sailing Notes
Notes
Food
Places
Croatia
North
Central
South
⛵ Sailing Food Plan & Shopping List
Duration: 6 Days | Crew: 8 People
📅 Daily Meal Schedule
| Day | Breakfast | Lunch / Salad | Main Dinner |
|---|---|---|---|
| Day 1 (Sun) | Daily offer / what’s found | Pasta with Roasted Pepper Pesto | Red Lentil Soup |
| Day 2 (Mon) | Daily offer / what’s found | Greek Baked Chickpeas | Harira Soup |
| Day 3 (Tue) | Daily offer / what’s found | Black Lentil Salad | Vegetable Risotto |
| Day 4 (Wed) | Daily offer / what’s found | Poke Bowl | Chili sin Carne |
| Day 5 (Thu) | Daily offer / what’s found | Shakshuka / Lečo | Butter Beans |
| Day 6 (Fri) | Daily offer / what’s found | Veggie Couscous | Leftover Feast |
See Shopping List for the full ingredient list and water requirements.
⚓ Yacht Cooking Tips
- Freshness Order: Use mushrooms, spinach, and soft tomatoes first. Cabbage, carrots, and potatoes can wait until the end of the week.
- Bread Management: Start with fresh bakery bread. Switch to vacuum-packed “bake-off” bread that you can toast in the oven from Day 4 onwards.
- Storage: Heavy items (water, cans, potatoes) should be stored low in the boat (the bilge) to keep the boat stable.
- Cleaning: Save fresh water by rinsing dishes in seawater first, then doing a final quick rinse with fresh water.
🍲 Recipes
- Pasta with Roasted Pepper Pesto
- Red Lentil Soup
- Greek Baked Chickpeas
- Harira Soup
- Black Lentil Salad
- Vegetable Risotto
- Poke Bowl
- Chili sin Carne
- Shakshuka / Lečo
- Butter Beans
- Veggie Couscous
- Porrusalda
🛒 Master Shopping List
Calculated for 8 People | 6 Days
🥬 Fresh Produce
| Item | Quantity | Used in |
|---|---|---|
| Onions (yellow) | ~15 pcs (~3 kg) | Red Lentil Soup, Greek Baked Chickpeas, Harira, Risotto, Chili, Shakshuka, Couscous |
| Onion (red) | 6 | Black Lentil Salad |
| Garlic | 5 heads | All savoury dishes |
| Leeks | 7 large stalks | Red Lentil Soup, Greek Baked Chickpeas |
| Carrots | ~2 kg (~15 pcs) | Red Lentil Soup, Harira, Risotto, Poke Bowl |
| Bell peppers | 20 pcs (mix of colors) | Chili sin Carne, Shakshuka |
| Zucchini | 6 | Vegetable Risotto, Veggie Couscous |
| Fresh tomatoes | 14 | Veggie Couscous (4), Butter Beans (4) |
| Cucumber | 8 | Poke Bowl |
| Avocado | 8 (mix ripe/firm) | Poke Bowl |
| Celery | 2 large bunches (~12 stalks) | Red Lentil Soup, Harira |
| Fresh parsley | 3 bunches | Black Lentil Salad, Shakshuka, Pasta with Tuna, Veggie Couscous |
| Fresh cilantro | 2 bunches | Harira, Veggie Couscous |
| Lemons | 6 | Red Lentil Soup, Harira |
🧀 Fridge & Proteins
| Item | Quantity | Used in |
|---|---|---|
| Eggs | 30 | Shakshuka + breakfasts |
| Klobása / hard salami | 400 g | Greek Baked Chickpeas |
| Tofu (firm) | 3 packs (~1.2 kg) | Greek Baked Chickpeas (400 g) + vegan Shakshuka option (800 g) |
| Parmesan | 400 g | Pasta with Pepper Pesto (80 g), Red Lentil Soup (rinds + garnish), Risotto (100 g) |
| Mozzarella | 4 balls | Black Lentil Salad |
| Smoked salmon | 800 g | Poke Bowl |
| Wakame (frozen) | 500 g | Poke Bowl |
| Hard cheese (Eidam / Gouda) | 1 kg | Breakfasts & snacking |
| Fresh spinach | 400 g | Butter Beans |
| Vegan heavy cream (oat/soy) | 4 packs (200 ml each) | Butter Beans |
🥫 Pantry & Dry Goods
| Item | Quantity | Used in |
|---|---|---|
| Pasta (penne / spaghetti) | 2 kg | Pasta with Pepper Pesto (1.5 kg) + extra |
| Arborio rice | 1 kg | Vegetable Risotto |
| Sushi rice | 1 kg | Poke Bowl |
| Couscous | 1 kg | Veggie Couscous |
| Buckwheat | 3 kg | Breakfasts |
| Vermicelli / soup noodles | 100 g | Harira |
| Canned crushed tomatoes | 10 cans (400 g) | Red Lentil Soup (2), Harira (2), Chili (2), Shakshuka (4) |
| Roasted red peppers (jars) | 4 jars (460 g each) | Pasta with Pepper Pesto |
| Tomato paste | 2 tubes (200 g each) | Harira, Chili, Veggie Couscous |
| Canned chickpeas | 8 cans (310 g each) | Greek Baked Chickpeas (4), Harira (4) |
| Canned beans (kidney or black) | 4 cans (400 g each) | Chili sin Carne |
| Canned corn | 6 cans (340 g each) | Vegetable Risotto (2), Chili sin Carne (2) |
| Canned peas | 6 cans (400 g each) | Vegetable Risotto (2), Veggie Couscous (2) |
| Canned mushrooms | 6 cans (400 g each) | Vegetable Risotto (2), Veggie Couscous (2) |
| Canned white beans | 8 cans (400 g each) | Butter Beans |
| Lentils (black / beluga) | 500 g | Black Lentil Salad |
| Lentils (red) | 1 kg | Red Lentil Soup (~500 g), Harira (200 g) |
| Olives (pitted) | 1 large jar (900 g) | Black Lentil Salad |
| Capers | 1 jar (100 g) | Black Lentil Salad |
| Sun-dried tomatoes (in oil) | 2 jars (280 g each) | Butter Beans |
| Nori / seaweed | 2 packs (25 g each) | Poke Bowl |
| Kimchi | 2 large jars (500 g each) | Poke Bowl |
| Edamame (canned) | 2 cans (400 g each) | Poke Bowl |
| Bamboo slices (canned) | 2 cans (540 g each, AROY-D) | Poke Bowl |
🧂 Spices & Essentials
| Item | Quantity | Used in |
|---|---|---|
| Olive oil | 2 L | All savoury dishes |
| Soy sauce | 1 bottle | Poke Bowl |
| Teriyaki sauce | 1 bottle | Poke Bowl |
| Rice vinegar | 1 bottle | Poke Bowl |
| Red wine vinegar | 1 bottle | Black Lentil Salad |
| Sriracha | 1 bottle | Poke Bowl |
| Wasabi | 1 small pack | Poke Bowl |
| Sesame seeds | 1 small pack | Poke Bowl |
| Harissa paste | 1 jar | Harira |
| Cumin | 1 jar | Harira, Chili, Shakshuka, Veggie Couscous |
| Smoked paprika | 1 jar | Harira, Chili, Shakshuka, Veggie Couscous |
| Cinnamon | 1 jar | Harira |
| Turmeric | 1 jar | Harira |
| Chili flakes | 1 jar | Chili sin Carne |
| Bay leaves | 1 pack | Red Lentil Soup |
| Dried sage or rosemary | 1 jar | Greek Baked Chickpeas |
| Dark chocolate | 1 bar | Chili sin Carne (1 square) |
| Sugar | 1 pack | Poke Bowl rice seasoning |
| Salt | 1 kg | All dishes |
| Black pepper | 1 grinder | All dishes |
| Vegetable bouillon | 3 packs (cubes/powder) | ~8 L broth across Red Lentil Soup, Harira, Risotto, Couscous |
| Bread | 12 loaves fresh (Days 1–3) + 6 packs bake-off baguettes (Days 4–6) | Serving with Greek Baked Chickpeas, Harira, Shakshuka |
| Snacks | 10 packs biscuits / crackers, 5 bags nuts, chips | — |
| Trash bags | 1 large pack | — |
| Baking paper | 1 roll | Greek Baked Chickpeas |
| Aluminium foil | 1 roll | Greek Baked Chickpeas |
🍺 Drinks
- Beer: 6 cans/person/day min
- Port wine
- Rum: 2 bottles min
- Coke + juice
💧 Water Requirements (3 L per person/day)
To ensure a safe supply for 8 people for 8 days (approx. 200 Liters):
- Option 1 (1.5 L Bottles): ~135 bottles (approx. 23 packs of six)
- Option 2 (2 L Bottles): ~100 bottles (approx. 17 packs of six)
Tip: Bring a permanent marker to label bottle caps with names to avoid waste.
Pasta with Roasted Pepper Pesto
Serves: 8 | Prep: 10 min | Cook: 15 min
Ingredients
- 1.5 kg pasta (penne or fusilli)
- 4 large jars roasted red peppers (or 8 fresh peppers, roasted and peeled)
- 100 ml olive oil
- 4 garlic cloves
- 80 g parmesan, grated
- Salt & pepper
Instructions
- Blend roasted peppers, olive oil, garlic, and parmesan until smooth. Season.
- Cook pasta al dente. Reserve 1 cup pasta water before draining.
- Toss pasta with pesto, loosening with pasta water as needed. Serve with extra parmesan.
Red Lentil Soup
Serves: 5 | Prep: 10 min | Cook: 50 min
Ingredients
- 4 tbsp extra virgin olive oil
- 2 medium yellow onions, medium dice
- 2 medium leeks, white and pale green parts only, rinsed and medium dice
- 3 medium carrots (~½ lb), medium dice
- 5 celery stalks (~6 oz), medium dice
- 2 large garlic cloves, roughly chopped
- 1½ cups red split lentils, rinsed
- 1 can (15½ oz) crushed Italian tomatoes
- 1–2 parmigiano-reggiano rinds
- 2 dried bay leaves
- 2 quarts (8 cups) low-sodium chicken or vegetable broth
- Kosher salt & freshly ground black pepper
- Freshly grated parmigiano-reggiano, for serving (optional)
Instructions
- Heat olive oil in a large soup pot over medium heat. Add onions, leeks, and a generous pinch of salt. Sauté with lid askew, stirring occasionally, until soft and translucent — about 10–15 minutes.
- Add carrots, celery, and garlic; stir for 3–4 minutes. Add lentils, crushed tomatoes, parmesan rinds, bay leaves, and broth. Bring to a boil, then reduce to medium-low and simmer uncovered for 30–40 minutes, stirring every 10 minutes, until vegetables are tender and lentils have broken down. Soup should be thick and hearty.
- Optionally blend a small portion with an immersion blender for a better texture.
- Season with salt and pepper. Add a squeeze of lemon if it tastes flat. Remove bay leaves and parmesan rinds before serving. Garnish with grated parmesan.
Greek Baked Chickpeas
Serves: 8 | Prep: 15 min | Cook: 60 min
Ingredients
- 4 cans chickpeas, drained
- 4 large leeks, white and pale green parts, sliced into rounds
- 2 onions, sliced
- 400 g tofu, cubed
- 400 g klobása / hard salami, sliced
- 6 tbsp olive oil
- 4 garlic cloves, minced
- 1 tbsp dried sage or rosemary
- 300 ml water or vegetable broth
- Salt & pepper
- Bread for serving
Instructions
- Preheat oven to 180°C.
- Combine chickpeas, leeks, onion, tofu, klobása, garlic, olive oil, and sage in a large baking dish. Add water and season generously.
- Cover with foil and bake for 45 minutes. Remove foil and bake a further 15–20 minutes until golden.
- Serve with crusty bread.
Harira Soup
Serves: 8 | Prep: 15 min | Cook: 40 min
Moroccan tomato soup with chickpeas and lentils — warming, spiced, and thick.
Ingredients
- 4 cans chickpeas, drained
- 200 g red lentils
- 2 onions, diced
- 4 celery stalks, diced
- 2 carrots, diced
- 2 cans crushed tomatoes
- 2 tbsp tomato paste
- 2 tbsp harissa paste
- 1 tsp turmeric
- 1 tsp cumin
- ½ tsp cinnamon
- 1 tsp smoked paprika
- 100 g vermicelli
- 2 L vegetable broth
- Olive oil, salt & pepper
- Fresh parsley & cilantro
- Lemon wedges for serving
Instructions
- Sauté onion, celery, and carrots in olive oil for 5 minutes until softened.
- Add garlic and spices; cook 1 minute.
- Add crushed tomatoes, tomato paste, lentils, chickpeas, and broth. Bring to a boil then simmer uncovered for 25 minutes.
- Add vermicelli and cook 8 minutes more. Stir in harissa.
- Season and serve with lemon wedges, fresh parsley, and bread.
Black Lentil Salad
Serves: 8 | Prep: 10 min | Cook: 25 min
Ingredients
- 500 g black / beluga lentils
- 1 red onion, finely diced
- 3 tbsp capers
- 200 g olives, pitted and halved
- 4 mozzarella balls, torn
- 4 tbsp olive oil
- 2 tbsp red wine vinegar
- Salt & pepper
- Fresh parsley
Instructions
- Cook lentils in salted boiling water for 20–25 minutes until tender. Drain and cool slightly.
- While still warm, toss with olive oil, vinegar, salt, and pepper.
- Mix in red onion, capers, and olives.
- Top with torn mozzarella and fresh parsley. Serve at room temperature.
Vegetable Risotto
Serves: 8 | Prep: 15 min | Cook: 35 min
Ingredients
- 1 kg arborio rice
- 1 kg mushrooms, sliced
- 2 zucchini, diced
- 2 cans peas, drained
- 2 cans corn, drained
- 2 carrots, diced
- 2 onions, diced
- 4 garlic cloves, minced
- 2 L vegetable broth, kept hot
- 100 g parmesan, grated
- Olive oil, salt & pepper
Instructions
- Sauté onion and garlic in olive oil until soft.
- Add mushrooms and carrots; cook until softened, about 8 minutes.
- Add rice and stir to coat. Pour in a ladle of hot broth; stir until absorbed. Repeat, adding broth one ladle at a time, for 18–20 minutes.
- Stir in zucchini, peas, and corn in the last 5 minutes.
- Finish with a drizzle of olive oil and parmesan. Season and serve immediately.
Poke Bowl
Serves: 8 | Prep: 20 min | Cook: 20 min
Ingredients
- 1 kg sushi rice
- 8 avocados, sliced
- 2 jars kimchi
- 2 packs nori / seaweed, cut into strips
- 4 cucumbers, sliced
- 4 carrots, julienned
- 4 tbsp soy sauce
- 2 tbsp sesame seeds
- Sriracha, wasabi & extra soy sauce for serving
Instructions
- Cook sushi rice per package instructions. Season with a splash of rice vinegar, a pinch of sugar, and salt.
- Divide rice into bowls.
- Arrange avocado, kimchi, seaweed, cucumber, and carrot on top.
- Drizzle with soy sauce and sprinkle sesame seeds. Serve with sriracha and wasabi on the side.
Chili sin Carne
Serves: 8 | Prep: 15 min | Cook: 35 min
Ingredients
- 4 cans kidney or black beans, drained
- 2 cans corn, drained
- 2 cans crushed tomatoes
- 4 bell peppers, diced
- 2 onions, diced
- 4 garlic cloves, minced
- 2 tbsp tomato paste
- 1 tbsp cumin
- 1 tbsp smoked paprika
- 1 tsp chili flakes
- 1 square dark chocolate
- Olive oil, salt & pepper
Instructions
- Sauté onion, peppers, and garlic in olive oil until soft, about 8 minutes.
- Add cumin, paprika, and chili flakes; cook 1 minute.
- Add beans, corn, crushed tomatoes, and tomato paste. Stir and simmer 25–30 minutes.
- Stir in dark chocolate at the end until melted. Season.
- Serve with bread or rice.
Shakshuka / Lečo
Serves: 8 | Prep: 10 min | Cook: 30 min
Eggs (or tofu) poached in spiced tomato and pepper sauce.
Ingredients
- 8 bell peppers, diced
- 4 cans crushed tomatoes
- 2 onions, diced
- 4 garlic cloves, minced
- 16 eggs (or 800 g firm tofu, cubed, for vegan)
- 2 tsp smoked paprika
- 1 tsp cumin
- Olive oil, salt & pepper
- Fresh parsley
Instructions
- Sauté onion and garlic in olive oil until soft.
- Add peppers; cook 5 minutes.
- Add crushed tomatoes and spices; simmer 15 minutes until sauce thickens.
- Make wells in the sauce and crack in eggs (or nestle tofu cubes). Cover and cook until eggs are just set, about 8–10 minutes.
- Season and garnish with parsley. Serve with bread.
Butter Beans
Serves: 8 | Prep: 10 min | Cook: 25 min
Creamy white bean stew with spinach, tomatoes, and sun-dried tomatoes in a rich vegan cream sauce.
Ingredients
- 8 cans white beans, drained
- 4 packs vegan heavy cream (oat or soy, 200 ml each)
- 400 g fresh spinach (or 500 g frozen)
- 2 jars sun-dried tomatoes in oil, roughly chopped
- 4 fresh tomatoes, diced
- 4 tbsp tomato puree
- 2 onions, diced
- 4 garlic cloves, minced
- Olive oil, salt & pepper
Instructions
- Sauté onion in olive oil over medium heat until soft, about 5 minutes. Add garlic and cook 1 minute.
- Add diced tomatoes, sun-dried tomatoes, and tomato puree. Simmer 8 minutes until tomatoes break down.
- Add white beans and vegan cream. Stir and simmer 10 minutes until sauce thickens.
- Fold in spinach and cook until just wilted, 2–3 minutes. Season with salt and pepper.
- Serve with crusty bread.
Veggie Couscous
Serves: 8 | Prep: 10 min | Cook: 20 min
Ingredients
- 1 kg couscous
- 2 zucchini, diced
- 2 cans peas, drained
- 4 tomatoes, diced
- 2 tbsp tomato paste
- 2 onions, diced
- 4 garlic cloves, minced
- 1 tsp cumin
- 1 tsp smoked paprika
- 2 tbsp olive oil
- ~1 L vegetable broth (hot, for couscous)
- Fresh parsley or cilantro, salt & pepper
Instructions
- Sauté onion and garlic in olive oil until soft.
- Add zucchini, cumin, and paprika; cook 5 minutes.
- Add diced tomatoes and tomato paste; simmer 10 minutes. Stir in peas. Season.
- Place couscous in a large bowl. Pour hot broth over in a 1:1 ratio, cover and rest 5 minutes, then fluff with a fork.
- Serve couscous topped with the vegetable sauce. Garnish with fresh parsley.
Porrusalda
Serves: 8 | Prep: 15 min | Cook: 55 min
Traditional Basque leek and potato soup — simple, hearty, and warming.
Ingredients
- 10 large leeks, white and pale green parts only, sliced into rounds and rinsed well
- 3 kg potatoes, broken into rough chunks (press with thumb against knife edge — don’t cut cleanly; rough edges release more starch)
- 3 carrots, diced
- 6 tbsp olive oil
- 4 garlic cloves, minced
- 2 bay leaves
- 2 L vegetable broth or water
- Salt & pepper
- Fresh parsley, crusty bread for serving
Instructions
- Sauté leeks in olive oil with a pinch of salt for 8–10 minutes until softened.
- Add garlic and carrots; cook 3 minutes.
- Add potato chunks, bay leaves, and broth. Bring to a boil.
- Reduce heat and simmer covered for 40–50 minutes until potatoes are very tender.
- Remove bay leaves. Optionally blend a small portion for a creamier texture.
- Season generously. Serve with crusty bread and a drizzle of olive oil.
Navigation — Staircase Method
The staircase method converts between four course types. The rule is simple:
- Going up (Compass → Water): ADD
- Going down (Water → Compass): SUBTRACT
The Staircase
┌──────────────────────────────────────────────────┐
│ KV — Water Course (Kurz vůči vodě) │
│ ↑ + drift ↓ − snos │
│ KR — True Course (Pravý kurz) │
│ ↑ + variation ↓ − variace │
│ KM — Magnetic Course (Magnetický kurz) │
│ ↑ + deviation ↓ − deviace │
│ KK — Compass Course (Kompasový kurz) │
└──────────────────────────────────────────────────┘
| Symbol | Name | Czech |
|---|---|---|
| KK | Compass Course | Kompasový kurz |
| KM | Magnetic Course | Magnetický kurz |
| KR | True Course | Pravý kurz |
| KV | Water Course | Kurz vůči vodě |
| dev | Deviation | Deviace |
| var | Variation | Variace |
| snos | Drift | Snos |
Signs
- W (West) var/dev → negative value
- E (East) var/dev → positive value
Variation (var)
Variation is the angle between True North and Magnetic North. Always update it from the chart for the current year:
Current var = Base var ± (annual change × years elapsed)
Example: Base 5°35’ W in 2018, annual change 6’ W, year 2023:
var = 5°35' + (5 × 6') = 5°35' + 30' = 6°05' W ≈ −6°
KK → KV (Compass to Water — going UP, ADD)
KM = KK + dev
KR = KM + var
KV = KR + snos
Example: KK = 261°, dev = 0°, var = −6°, wind from starboard (snos = −4°)
KM = 261° + 0° = 261°
KR = 261° + (−6°) = 255°
KV = 255° + (−4°) = 251°
KV → KK (Water to Compass — going DOWN, SUBTRACT)
KR = KV − snos
KM = KR − var
KK = KM − dev
Example: KR = 346°, dev = 0°, var = −6°
KM = 346° − (−6°) = 352°
KK = 352° − 0° = 352°
Drift (snos)
Drift is the sideways push of the wind on the boat, shifting the water course from the true course.
| Wind direction | drift sign |
|---|---|
| Wind from starboard (right) | − (subtract) |
| Wind from port (left) | + (add) |
Bearings
The same staircase applies when converting compass bearings to true bearings for chart plotting:
True bearing = Compass bearing + dev + var
Plot two or more true bearings on the chart — their intersection is your estimated position.
Distance & Speed
Distance (NM) = Speed (kts) × Time (min) / 60
Time (min) = Distance (NM) × 60 / Speed (kts)
Example: Speed = 3 kts, Time = 90 min
Distance = 3 × 90 / 60 = 4.5 NM
ETA
ETA = departure time + travel time
travel time (min) = Distance × 60 / Speed
Example: Depart 16:00, distance = 9.6 NM, speed = 3 kts
travel time = 9.6 × 60 / 3 = 192 min = 3h 12min
ETA = 16:00 + 3:12 = 19:12
Converting Compass Bearing to Chart Bearing (Position Fix)
When you sight a landmark with a hand bearing compass, the reading is a compass bearing. To plot it on a nautical chart (which uses true north) you must convert it using the same staircase.
Steps
- Sight the landmark with the hand bearing compass → get compass bearing
- Convert to true bearing (go UP, ADD):
True Bearing = Compass bearing + dev + var - Plot on chart: from the landmark, draw a line in the reciprocal direction (true bearing ± 180°) — you are somewhere along this line
- Repeat for a second (and ideally third) landmark
- Intersection of the lines = your position fix
Example
Compass bearing to lighthouse = 043°, dev = 0°, var = −6°
True Bearing = 043° + 0° + (−6°) = 037°
Reciprocal = 037° + 180° = 217°
Draw a line from the lighthouse at 217° on the chart. You are on that line.
Tips
- Use 3 bearings when possible — if they form a small triangle (cocked hat), your position is inside it
- Choose landmarks roughly 60–120° apart for the best intersection angle
- Bearings to objects ahead or astern are less reliable — prefer objects to the side
- Take all bearings quickly to minimise error from boat movement
Measuring Compass Deviation
Deviation is caused by the boat’s own magnetic field (engine, wiring, steel fittings). It varies with heading and must be measured for each compass on each boat.
Formula
dev = True Course − var − KK
Where True Course comes from GPS, a transit, or a known bearing.
Method 1: GPS Comparison (easiest)
Modern GPS units show COG (Course Over Ground) in true degrees.
Requirements: calm water, no current, no wind (so the boat tracks straight — leeway = 0).
Steps:
- Motor (not sail — sails cause leeway) on a steady heading
- Note the compass heading (KK)
- Read the GPS COG (= True Course)
- Calculate deviation:
dev = GPS COG − var − KK - Repeat on at least 8 headings (N, NE, E, SE, S, SW, W, NW)
- Record results in a deviation table
Example: KK = 090°, var = −6°, GPS COG = 087°
dev = 087° − (−6°) − 090° = 087° + 6° − 090° = +3°E
Method 2: Transit / Leading Lines (most accurate)
A transit is two charted objects that line up — their true bearing is known from the chart.
Steps:
- Find two charted objects that form a transit (e.g. lighthouse and church spire)
- Look up or calculate the true bearing of the transit from the chart
- Motor until both objects are exactly in line
- Note the compass heading at that moment
- Calculate:
dev = True Bearing − var − Compass reading
Advantage: No GPS needed, very precise.
Method 3: Reciprocal Bearings
Use a hand bearing compass (minimally affected by ship’s magnetism) as a reference.
Steps:
- Anchor or stop the boat
- Send a person ashore with the hand bearing compass
- Both take bearings to each other simultaneously
- The two bearings should differ by exactly 180°
- Any difference is deviation in the steering compass
Deviation Table
After measuring on multiple headings, record the results:
| Ship’s Head (KK) | Deviation |
|---|---|
| 000° (N) | — |
| 045° (NE) | — |
| 090° (E) | — |
| 135° (SE) | — |
| 180° (S) | — |
| 225° (SW) | — |
| 270° (W) | — |
| 315° (NW) | — |
Interpolate between measured headings for any course in between.
Tips
- Deviation changes if you add/move metal objects, electronics, or speakers near the compass
- Remeasure after any significant modification to the boat’s equipment
- A deviation of less than ±3° is generally acceptable for coastal sailing
- Steer clear of metal objects and active electronics when taking compass readings
Quick Reference
| Goal | Formula | Rule |
|---|---|---|
| KK → KM | KM = KK + dev | Go UP, ADD |
| KM → KR | KR = KM + var | Go UP, ADD |
| KR → KV | KV = KR + snos | Starboard: −, Port: + |
| KV → KR | KR = KV − snos | Go DOWN, SUBTRACT |
| KR → KM | KM = KR − var | Go DOWN, SUBTRACT |
| KM → KK | KK = KM − dev | Go DOWN, SUBTRACT |
| Compass → Chart bearing | TB = KB + dev + var | Go UP, ADD |
| Plot position line | Reciprocal = TB ± 180° | Draw from landmark |
| Distance | D = V × t(min) / 60 | — |
| Time | t = D × 60 / V | — |
Sailing Logbook
Totals:
- As skipper: ~848NM
- Totals: ~948NM
Season 2026
- Distance: TBA NM
Season 2025
- Distance: ~105 NM
Season 2023
- Distance: 265 NM
- Maximum speed: 5.9 kts
- Average speed: 3.9 kts
Season 2022
- Distance: 208 NM
- Maximum speed: 7.6 kts
- Average speed: 3.1 kts
Season 2021
- Distance: 183 NM
- Maximum speed: 9.9 kts
- Average speed: 4.3 kts
Season 2020
- Distance: 87 NM
- Maximum speed: 7.1 kts
- Average speed: 3.1 kts
Season 2019
- Distance 100NM [^1]
[^1] 100NM sailed in baltic sea as training
2026
Marina Sibenik 30.05.2026 - 06.06.2026
Trip summary:
- Distance: TBA NM
- Maximum speed: - kts
- Average speed: - kts
- Boat: Dufour 460 GL | Almar
- Year: 2017
- Drought: 2.20 m
- Length: 14.15 m
- Beam: 4.50 m
- Engine: 75 hp (55,16 kW)
- Fuel tank: 250 l
- Water tank: 530 l
- Classical mainsail
Map

Route stops and details:
Trip Gallery
Planning
Sailing Notes
Generic Check
- Passport / ID card
- Driving licence
- Boat booking confirmation
- Insurance
- Crew list
- Skipper licence
- Boat contract / charter agreement
- Cash
- Payment cards
- Nautical charts & manual navigation tools
Personal Checklist
Clothing
- Hat/Cap
- Sunglasses + strap
- UV protective long-sleeve shirt (sun on open water is intense)
- Short-sleeve shirts
- Shorts
- Long trousers (evenings, March-April)
- Swimwear
- Light fleece / hoodie (cool nights, early season)
- Windbreaker / windstopper
- Light rain jacket (squalls occur even in summer)
- Shoes for land + sandals / flip flops
- Light-soled shoes or deck boots for the boat
- Reef shoes / water shoes (rocky Med beaches)
- Sailing gloves
Gear
- Travel pillow
- Headlight (with red light)
- Snorkeling mask + fins
- Binoculars
- Powerbank
- Hammock
Hygiene & Health
- Towel (quick-dry)
- Sunscreen SPF 50+
- After-sun lotion
- Lip balm with SPF
- Insect repellent (mosquitoes are common in marinas at dusk)
- Rehydration salts / electrolytes (heat exhaustion risk)
- Antihistamines (fenistil — jellyfish stings, insect bites)
- Eye drops (sun and salt air)
- Soap / shower gel
- Medicines
- Nausea pills (ginger sweets / lollipops)
Entertainment & Extras
- Musical instruments / songbooks
- Books
Boat Supplies
Kitchen & Cooking
- Aeropress / moka pot / French press
- Thermos flask
Tools & Repair
- Rope / string (1m short + longer line)
- Matches x3 + lighter + candles
- Needle & thread + plasters
- Duck tape
- Carpet tape
- Screws
- Phillips screwdriver
- Soldering gun
Pre-departure Checklist
- Knots practice: cleat hitch, clove hitch, figure-eight, bollard tie, bowline
- Boat orientation:
- Lines in lockers — which does what
- Mainsail halyard
- Furling jib
- Main sheet & jib sheet
- Winches
- Anchor & controller
- Autopilot is used only on engine or when waves are calm and the boat is well balanced from a sail trim perspective
- Prepare helm for quick deployment
- Always keep everything secured — the sea is not always calm and everything follows gravity
- Install safety lines and safety net on railing (when Amelia is on board)
Rules
- The designated skipper is always right — he is currently responsible for safety and navigation
- Skipper is responsible for safety — follow instructions immediately, without discussion
- Do not stand on cabin steps
Handover inspection
- Rudder
- Boom joint
- Furling jib lift
- Aft leech tensioner
- Navigational lights (port, starboard, masthead, stern, anchor light)
- Navigational ball and triangle (day shapes)
- Anchor (condition, chain, windlass operation)
- Autopilot (function, drive unit)
- Speedometer / log
- GPS / chartplotter
- Fuses (location, spares)
- Tools (location, contents)
- Engine room (belts, hoses, bilge, seacocks)
- Oil level — engine (dipstick fully inserted to read) and saildrive (dipstick rested on thread, do not screw in)
- Toilets & black water tank (check integrity with a bit of saponate — look for foam or food coloring at through-hulls)
Adriatic Winds
| Wind | Direction | Season | Typical Force | Character | Hazards & Notes |
|---|---|---|---|---|---|
| Bura | NE – ENE | Oct – Apr; possible year-round | 6 – 10+ (gusts >60 kn) | Cold, dry, katabatic; descends from Dinaric Alps; extremely gusty | Most dangerous Adriatic wind; strongest in Kvarner, Velebit channel, and Trieste; can appear with little warning; short steep seas |
| Jugo | SE – S | Autumn, winter, spring | 5 – 8 | Warm, humid, persistent; builds slowly over 1–3 days; long-period swell | Poor visibility, rain; seas become very rough and confused; subsides slowly; watch for prolonged forecasts |
| Maestral | NW – WNW | Jun – Sep (afternoons) | 3 – 5 | Thermal sea breeze; develops late morning, peaks early afternoon, dies at sunset | Reliable summer sailing wind; pleasant conditions; no significant swell; can build to F5–6 on exposed coasts |
| Tramontana | N – NNE | Autumn, spring | 4 – 7 | Cold, dry, continental; steadier and less gusty than Bura | Brings clear skies; chilly; can be confused with Bura onset |
| Nevera | NW – N (squall) | Jun – Sep | Squall: 7 – 10+ | Sudden violent thundersquall; forms rapidly over mountains in summer heat | Most dangerous in summer; little warning; waterspouts possible; seek shelter immediately when cumulonimbus build inland |
| Lebić | SW – WSW | Autumn, winter | 4 – 7 | Warm, humid; brings swell from the open Mediterranean | Rain and poor visibility; less common than Jugo; can combine with Jugo to create confused cross-seas |
| Ostro | S – SSE | Variable | 3 – 6 | Warm, humid southerly; precursor or variant of Jugo | Brings mist and cloud; often transitions into full Jugo |
| Burin | NE – variable (offshore) | Summer nights | 1 – 3 | Light nighttime land breeze; calm, dry; fades at sunrise | Complement to daytime sea breeze; useful for motoring out of anchorage at night |
| Pulenat | W | Variable | 3 – 5 | Moderate westerly; relatively rare in central Adriatic | Can bring short choppy seas on exposed passages |
| Levanat | E – ENE | Variable | 3 – 5 | Easterly; uncommon in central and southern Adriatic | May bring spray and choppy conditions; more frequent in northern Adriatic |
| Grego | Greco | NE (between Tramontana and Levanat) | Autumn, winter | 4 – 7 | Cold and dry; similar to Tramontana; can be gusty near capes |
| Vardar | — | NE – E | Winter | 5 – 8 | Cold, dry katabatic wind funnelled through river valleys in Greece/N Macedonia; reaches southern Adriatic |
2025
Marina Preveza 14.06.2025 - 28.06.2025
Trip summary:
- Distance: ~105 NM
- Maximum speed: - kts
- Average speed: - kts
- Boat: Bavaria Cruiser 41 | Anemoessa
- Year: 2016
- Drought: 1.85 m
- Length: 13.10 m
- Beam: 3.99 m
- Engine: 55 hp
- Fuel tank: 210 l
- Water tank: 360 l
- Furling mainsail
Map

Route stops and details:
Trip Gallery
2023
Marina Veruda 29.04.2023 - 06.05.2023
Trip summary:
- Distance: 215 NM
- Maximum speed: 6.08 kts
- Average speed: 3.88 kts
- Boat: Elan 40.1 Impression | Estela
- Year: 2020
- Drought: 1.80 m
- Length: 11.83 m
- Beam: 3.91 m
- Engine: 40 hp (29.8 kW)
- Fuel tank: 146 l
- Water tank: 400 l
- Classic mainsail
Map

Route stops and details:
Trip Gallery
Marina Punat 16.09.2023 - 23.09.2023
Trip summary:
- Distance: ~50 NM
- Maximum speed: - kts
- Average speed: - kts
- Boat: Dufour 430 | Catacea
- Year: 2020
- Drought: 2.10 m
- Length: 13.23 m
- Beam: 4.30 m
- Engine: 75 hp (55.9 kW)
- Fuel tank: 250 l
- Water tank: 530 l
- Classic mainsail
Map

Route stops and details:
Trip Gallery
2022
Marina Pomer 07.05.2022 - 14.05.2022
Trip summary:
- Distance: 208 NM
- Maximum speed: 7.5 kts
- Average speed: 3.3 kts
- Boat: Beneteau Oceanis 35 | DalMar
- Year: 2016
- Drought: 1.85m
- Length: 9.99 m
- Beam: 3.7 m
- Engine: 29 hp (21.6 kW)
- Fuel tank: 130 l
- Water tank: 330 l
- Classic mainsail
Map

Route stops and details:
- Start: Marina Pomer ↓
- End: Marina Pomer
- GPX
Trip Gallery
2021
Marina Frapa Rogoznica 29.05.2021 - 05.06.2021
Trip summary:
- Distance: 183 NM
- Maximum speed: 9.9 kts
- Average speed: 4.3 kts
- Boat: Dufour 430 | Calando
- Year: 2020
- Drought: 2.10m
- Length: 13.24 m
- Beam: 4.3 m
- Engine: 60 hp (44.7 kW)
- Fuel tank: 200 l
- Water tank: 380 l
- Rolling mainsail
Map

Route stops and details:
- Start: Marina Frapa ↓
- End: Marina Frapa
- GPX
Trip Gallery
2020
Vodice 05.09.2020 - 12.09.2020
Trip summary:
- Distance: 87 NM
- Maximum speed: 7.1 kts
- Average speed: 3.1 kts
- Boat: Jeanneau Sun Odyssey 30 | Espresso 1
- Year: 2009
- Drought: 1.95m
- Length: 8.99 m
- Beam: 3.18 m
- Engine: 21 hp (15.7 kW)
- Fuel tank: 50 l
- Water tank: 160 l
- Rolling mainsail
Route details
Trip Gallery
Sailing Galleries
Veruda - Dragove - Rovijn - Veruda 2023
Pomer - Veli Rat - Pomer 2022
Frapa - Vis - Frapa 2021
Vodice - Zut - Vodice 2020
Filesystem
XFS
XFS logging
Overview
XFS uses Write-Ahead Logging (WAL) to guarantee filesystem metadata consistency. Every metadata change is recorded in the log before being applied to the on-disk structures. If a crash occurs, the log is replayed to bring the filesystem back to a consistent state.
The logging subsystem has two major layers that work together:
- The circular on-disk log — a fixed-size ring buffer of 512-byte blocks stored in a dedicated log device (or the end of the data device).
- The Committed Item List (CIL) / Delayed Logging layer — an in-memory aggregation layer that batches and de-duplicates log writes before flushing to the on-disk log.
Key Data Structures
struct xlog — The Log Manager
Defined in xfs_log_priv.h, this is the central control structure for the entire
logging subsystem.
struct xlog {
struct xfs_mount *l_mp; // owning filesystem mount
struct xfs_ail *l_ailp; // Active Item List
struct xfs_cil *l_cilp; // Committed Item List (delayed logging)
struct xlog_grant_head l_reserve_head; // logical reservation accounting
struct xlog_grant_head l_write_head; // physical space accounting
atomic64_t l_tail_lsn; // LSN of oldest unpersisted transaction
struct xlog_in_core *l_iclog; // head of the iclog ring
spinlock_t l_icloglock; // protects the iclog state machine
};
The two grant_head fields are the heart of log space management and are explained
in detail in Log Space Accounting.
struct xlog_in_core — The In-Core Log Buffer (iclog)
Each iclog is a chunk of memory that absorbs formatted log records before they are written to disk. They form a circular ring buffer of typically 4 to 8 buffers, each up to 256 KB.
struct xlog_in_core {
enum xlog_iclog_state ic_state; // state machine position
atomic_t ic_refcnt; // reference count
wait_queue_head_t ic_force_wait; // waiters on forced flush
struct xlog_in_core *ic_next; // next in ring
u32 ic_offset; // current write cursor
u32 ic_size; // total buffer size
void *ic_datap; // pointer to buffer data
struct list_head ic_callbacks; // CIL checkpoint callbacks
};
Iclog state machine (xfs_log.c):
ACTIVE → WANT_SYNC → SYNCING → DONE_SYNC → CALLBACK → DIRTY → ACTIVE
| State | Meaning |
|---|---|
ACTIVE | Accepting new log records |
WANT_SYNC | Full or flushed, waiting for all writers to finish |
SYNCING | I/O submitted to disk |
DONE_SYNC | I/O complete, callbacks pending |
CALLBACK | Running CIL checkpoint callbacks (AIL insertion) |
DIRTY | Buffer spent; being recycled back to ACTIVE |
struct xfs_cil and struct xfs_cil_ctx — The Delayed Logging Layer
xfs_cil_ctx (xfs_log_priv.h) is the container for a single CIL checkpoint —
a batch of log items accumulated since the last checkpoint flush.
struct xfs_cil_ctx {
xfs_csn_t sequence; // monotonically increasing checkpoint number
xfs_lsn_t start_lsn; // LSN of first log record
xfs_lsn_t commit_lsn; // LSN of commit record
struct list_head lv_chain; // chain of formatted shadow buffers
atomic_t space_used; // bytes accumulated so far
struct xlog_ticket *ticket; // log reservation ticket
};
xfs_cil (xfs_log_priv.h) owns the current context and drives the push:
struct xfs_cil {
struct xlog *xc_log;
struct rw_semaphore xc_ctx_lock; // write-locked during push only
struct xfs_cil_ctx *xc_ctx; // current live context
spinlock_t xc_push_lock; // protects ordering list
wait_queue_head_t xc_commit_wait;
void __percpu *xc_pcp; // per-CPU item lists
};
struct xfs_log_vec — The Shadow Buffer
When a transaction commits, each log item’s in-memory state is formatted into a
shadow buffer (a log_vec) and decoupled from the live object. This is the
central innovation of delayed logging.
struct xfs_log_vec {
struct list_head lv_list; // CIL chain
uint32_t lv_order_id; // intra-checkpoint ordering
int lv_niovecs; // number of iovecs
struct xfs_log_iovec *lv_iovecp; // formatted region descriptors
struct xfs_log_item *lv_item; // back-pointer to log item
char *lv_buf; // shadow buffer memory
int lv_bytes; // bytes used
};
After formatting, the log item is unlocked immediately. The shadow buffer holds all data required for the eventual log write.
Log Space Accounting: The Dual-Grant-Head Model
XFS tracks log space with two independent accounting heads, both defined as
xlog_grant_head (xfs_log_priv.h):
| Head | What it tracks | Can overcommit? |
|---|---|---|
l_reserve_head | Logical reservation (space promised to transactions) | Yes — future commits block, but existing ones proceed |
l_write_head | Physical bytes actually written to the log | No — hard limit, never advances past the tail |
Why two heads? Rolling transactions (e.g., directory operations that span many buffer modifications) need to reserve space upfront without exhausting the physical log. The reserve head allows logical overcommitment so a rolling transaction can keep rolling, while the write head enforces the actual circular buffer boundary.
CIL space limits (xfs_log_priv.h):
// Background push triggered at ~12.5% of total log size
#define XLOG_CIL_SPACE_LIMIT(log) min_t(int, (log)->l_logsize >> 3, ...)
// Transaction commits throttled (sleeping wait) at 25% of log size
#define XLOG_CIL_BLOCKING_SPACE_LIMIT(log) (XLOG_CIL_SPACE_LIMIT(log) * 2)
LSN Encoding
Log Sequence Numbers are 64-bit values:
LSN = (cycle << 32) | block_offset
- cycle: how many times the log has wrapped around.
- block_offset: 512-byte block offset within the log.
Macros CYCLE_LSN() and BLOCK_LSN() extract these fields throughout the code.
Transaction Lifecycle with Delayed Logging
This is the full path from a filesystem operation to a durable on-disk record.
Phase 1: Transaction Allocation and Reservation
xfs_trans_alloc()
→ xlog_ticket_alloc()
→ xlog_grant_head_check() ← may sleep here if log is full
A ticket (reservation) is allocated holding the worst-case byte count for this transaction type. The reservation is calculated at mount time from geometric properties (tree depth, block size, etc.) and accounts for recursive modifications.
Phase 2: Item Modification
The caller modifies in-memory metadata (inode, buffer, dquot). Items are logged
via xfs_trans_log_inode(), xfs_trans_log_buf(), etc., which mark items dirty on
the transaction.
Phase 3: Transaction Commit → CIL Insertion
xfs_trans_commit() → xlog_cil_commit():
- Shadow buffer allocation (
xlog_cil_alloc_shadow_bufs()): outsidexc_ctx_lock, allocate memory sized for each item’s formatted representation. - Lock acquisition: acquire
xc_ctx_lockas a reader (allows concurrent commits). - Format into shadow buffers: call each item’s
iop_format()callback, writing item state into the shadow buffer. - Pin on first insertion: if the item has not been in the CIL before, call
iop_pin(). Re-logging (subsequent modifications before a checkpoint) does not add additional pins — the existing pin is reused and the old shadow buffer is discarded. - Add to per-CPU list: attach the
log_vecto the per-CPU CIL pending list. - Unlock item: the live object is immediately available for further modification.
- Release
xc_ctx_lock.
Relogging
Relogging is the critical property that prevents log tail pinning. Each re-commit of
an item supersedes the previous version. Only the latest aggregate snapshot is
written to the on-disk log (xfs-delayed-logging-design.rst):
Transaction What is logged LSN
A A X
B A+B X+n ← A's slot at X is now stale
C A+B+C X+n+m
Phase 4: CIL Background Push
The CIL push worker (xlog_cil_push_work(), xfs_log_cil.c) is triggered when:
- CIL space exceeds
XLOG_CIL_SPACE_LIMIT(background push), or - A caller explicitly issues
xfs_log_force()orfsync().
Push sequence:
1. Acquire xc_ctx_lock as WRITER (excludes all new commits)
2. Swap context: xc_ctx points to a new empty context
3. Release xc_ctx_lock ← concurrent commits resume on new context
4. Sort and aggregate log_vecs from the old context
5. Write start record and all item vectors to iclogs
6. Order commit record among concurrent checkpoints (xlog_cil_order_write)
7. Write commit record to iclog; submit iclog for I/O
8. On I/O completion: insert items into AIL, call iop_committed, unpin items
Step 6 — checkpoint ordering — ensures commit records appear in the log in strictly ascending checkpoint sequence order, regardless of when individual iclogs complete I/O.
Phase 5: Iclog I/O and AIL Insertion
When an iclog transitions from SYNCING to DONE_SYNC:
xlog_state_iodone_process_iclog()runs callbacks.xlog_cil_process_committed()inserts log items into the Active Item List (AIL) at their commit LSN.- Items are unpinned (
iop_unpin()), making them eligible for writeback.
Phase 6: AIL Writeback and Log Tail Advancement
Once items are inserted into the AIL the xfsaild kernel thread takes over. The
full mechanism is described in The Active Item List and Log Tail Pushing.
Log space is only reclaimed when the tail advances. This makes the log a true circular buffer.
The Active Item List and Log Tail Pushing
The AIL is the bridge between the log and the on-disk metadata. Its sole purpose is to track every log item that has been committed to the log but not yet written to its final on-disk location, and to push those items to disk so the log tail can advance and log space can be reclaimed.
Data Structures
struct xfs_ail (xfs_trans_priv.h)
struct xfs_ail {
struct xlog *ail_log; // log being managed
struct task_struct *ail_task; // xfsaild kthread
struct list_head ail_head; // LSN-ordered item list
struct list_head ail_cursors; // active traversal cursors
spinlock_t ail_lock; // protects all AIL state
xfs_lsn_t ail_last_pushed_lsn;// LSN of last successfully pushed item
xfs_lsn_t ail_head_lsn; // log head LSN at AIL init
int ail_log_flush; // counter: force CIL push when set
unsigned long ail_opstate; // XFS_AIL_OPSTATE_PUSH_ALL flag
struct list_head ail_buf_list; // buffers queued for delwri submission
wait_queue_head_t ail_empty; // waiters for AIL to drain completely
xfs_lsn_t ail_target; // LSN we are currently pushing toward
};
The list at ail_head is kept in strict ascending LSN order. The item at the front
(minimum LSN) defines the log tail: that is the oldest record in the log that has
not yet been written to disk.
struct xfs_ail_cursor (xfs_trans_priv.h)
struct xfs_ail_cursor {
struct list_head list; // registered in ailp->ail_cursors
struct xfs_log_item *item; // current position (low bit = invalidated)
};
Cursors allow xfsaild to walk the AIL safely even when items are deleted
concurrently. When an item is removed, every cursor pointing at it has its low
pointer bit set. The next call to xfs_trans_ail_cursor_next() detects this and
restarts the traversal from the new minimum.
The xfsaild Daemon Main Loop (xfs_trans_ail.c:653)
xfsaild is a single per-filesystem kthread. Its loop has three states:
┌──────────────────────────────────────────────────────┐
│ Set TASK_KILLABLE or TASK_INTERRUPTIBLE │
│ (KILLABLE if tout ≤ 20ms for fast wakeup) │
├──────────────────────────────────────────────────────┤
│ Check kthread_should_stop() → drain ail_buf_list │
│ and exit on shutdown │
├──────────────────────────────────────────────────────┤
│ If AIL empty AND ail_buf_list empty → schedule() │
│ (full idle: no timeout, wait for wakeup) │
├──────────────────────────────────────────────────────┤
│ If tout > 0 → msleep(tout) │
├──────────────────────────────────────────────────────┤
│ tout = xfsaild_push(ailp) ← core work │
└──────────────────────────────────────────────────────┘
The thread is woken by:
xfs_ail_push()— called fromxlog_assign_tail_lsn()when the log approaches full.xfs_ail_push_all()— called by umount and log quiesce to drain the AIL completely.xfs_trans_ail_update_bulk()— any new insertion into the AIL.
Push Target Calculation (xfs_trans_ail.c:405)
Before scanning items, xfsaild_push() calls xfs_ail_calc_push_target() to
decide how far to push. The logic in order of priority:
-
Push-all flag set (
XFS_AIL_OPSTATE_PUSH_ALL) orail_emptyhas waiters: returnmax_lsn(the current log head). Push everything. -
Log already has ≥ 25% free space:
free_bytes = l_logsize − (head_lsn − min_lsn) if free_bytes ≥ l_logsize / 4 → keep current ail_targetNo pushing needed; keep the existing target.
-
Log has < 25% free space: advance the target by 25% of the log size from the current tail:
target_block = BLOCK_LSN(min_lsn) + (l_logBBsize >> 2); // wrap cycle if needed target_lsn = xlog_assign_lsn(target_cycle, target_block);The target is clamped to
max_lsnand never lowered below the existingail_target.
Design intent: one push round reclaims exactly 25% of the log, ensuring a predictable amount of free space without over-flushing.
The xfsaild_push() Loop (xfs_trans_ail.c:535)
This is the core of the push. It runs under ail_lock for the traversal, briefly
dropping it during I/O operations.
Step 1: CIL Pre-flush Optimization
if (ailp->ail_log_flush && ailp->ail_last_pushed_lsn == 0 &&
(!list_empty_careful(&ailp->ail_buf_list) || xfs_ail_min_lsn(ailp))) {
ailp->ail_log_flush = 0;
xlog_cil_flush(ailp->ail_log);
}
When the AIL has items but the push cursor is stuck at the beginning
(ail_last_pushed_lsn == 0), it means items are pinned by in-flight CIL
transactions. Rather than spinning on pinned items, xfsaild forces a synchronous
CIL flush. This breaks the potential circular wait:
CIL holds pins → AIL cannot advance → log fills → CIL cannot commit → deadlock
Step 2: Cursor Initialization and Target Update
WRITE_ONCE(ailp->ail_target, xfs_ail_calc_push_target(ailp));
lip = xfs_trans_ail_cursor_first(ailp, &cur, ailp->ail_last_pushed_lsn);
The cursor starts from ail_last_pushed_lsn so that a push that hit the item limit
in one round can continue from where it left off in the next.
Step 3: Item Traversal
while (XFS_LSN_CMP(lip->li_lsn, ailp->ail_target) <= 0) {
if (test_bit(XFS_LI_FLUSHING, &lip->li_flags))
goto next_item; // skip: already in-flight
xfsaild_process_logitem(ailp, lip, &stuck, &flushing);
count++;
if (stuck > 100)
break; // backoff: too many blocked items
if (lip->li_lsn != lsn && count > 1000)
break; // per-LSN limit: avoid infinite loop
}
Two hard limits prevent the push loop from monopolizing the CPU:
stuck > 100: if more than 100 consecutive items are pinned or locked, abort and sleep. Continuing would just burn CPU with no progress.count > 1000at a new LSN: prevents unbounded iteration when many items share the same commit LSN.
Step 4: Async Buffer Submission
if (xfs_buf_delwri_submit_nowait(&ailp->ail_buf_list))
ailp->ail_log_flush++;
All buffers queued during the traversal are submitted in a single batched write.
submit_nowait returns non-zero if the submission was not possible (e.g. I/O error
or congestion), which sets ail_log_flush to trigger a CIL flush on the next round.
Step 5: Timeout Selection
The return value controls how long xfsaild sleeps before the next round:
| Condition | tout | Meaning |
|---|---|---|
| Reached target, or AIL empty | 50 ms | Wait for in-flight I/O to complete; reset cursor to 0 |
| >90% of items were stuck/flushing | 20 ms | Back off; next round may issue a log force; reset cursor to 0 |
| More items remain below target | 0 ms | Return immediately; continue from ail_last_pushed_lsn |
The cursor reset (ail_last_pushed_lsn = 0) on the first two cases ensures the
next wakeup re-evaluates the entire AIL from the minimum, picking up items that may
have been unpinned during the sleep.
Per-Item Push: xfsaild_process_logitem() (xfs_trans_ail.c:468)
For each item in the traversal, xfsaild_push_item() dispatches to the item’s
iop_push callback and interprets the return code:
| Return code | Meaning | Action |
|---|---|---|
XFS_ITEM_SUCCESS | Queued for I/O | Update ail_last_pushed_lsn |
XFS_ITEM_FLUSHING | Already being written | Increment flushing; update ail_last_pushed_lsn |
XFS_ITEM_PINNED | Held by an in-memory transaction | Increment stuck; set ail_log_flush |
XFS_ITEM_LOCKED | Could not acquire buffer lock | Increment stuck |
XFS_ITEM_FAILED | Previous I/O failed | Resubmit via xfsaild_resubmit_item() |
Items with XFS_LI_FAILED set are handled by xfsaild_resubmit_item() which
re-queues the backing buffer directly to ail_buf_list without calling iop_push
again, allowing the I/O to be retried on the next submission round.
Inode Item Push: xfs_inode_item_push() (xfs_inode_item.c:739)
Inode items use cluster flushing to amortize I/O overhead. A cluster is a group of inodes that share a single filesystem buffer (typically a 4 KB or 16 KB block).
1. Check preconditions (return PINNED or FLUSHING if not ready):
- inode stale (being freed)? → PINNED
- ipincount > 0? → PINNED
- cluster buffer pinned? → PINNED
- XFS_IFLUSHING flag set? → FLUSHING
- xfs_buf_trylock() fails? → LOCKED
2. Release ail_lock ← avoids holding spinlock during I/O
3. xfs_iflush_cluster(bp) ← formats ALL inodes in the cluster into bp
4. xfs_buf_delwri_queue(bp, &ailp->ail_buf_list)
5. Reacquire ail_lock
xfs_iflush_cluster() walks all inodes mapped to the same buffer and formats each
one’s in-memory xfs_dinode into the buffer in one pass. This means that when
xfsaild pushes one inode item, it potentially writes dozens of inodes with a
single I/O, which is critical for performance on inode-dense workloads.
The AIL lock is dropped during xfs_iflush_cluster(). Cursors handle any
concurrent deletions that occur during this window.
Buffer Item Push: xfs_buf_item_push() (xfs_buf_item.c:565)
Buffer items (btree blocks, superblock, AGF/AGI headers, etc.) have a simpler push path:
1. xfs_buf_ispinned(bp)? → PINNED (transaction holds a log reference)
2. xfs_buf_trylock(bp) fails?→ LOCKED (re-check pin after trylock failure)
3. Log a warning if XBF_WRITE_FAIL is set (previous write error)
4. xfs_buf_delwri_queue(bp, &ailp->ail_buf_list)
5. xfs_buf_unlock(bp)
Unlike inodes, each buffer item maps 1:1 to a buffer, so no clustering is needed.
The trylock avoids blocking — if the buffer is locked by another writer, xfsaild
moves on and returns to it on the next round.
Tail Advancement: __xfs_ail_assign_tail_lsn() (xfs_trans_ail.c:753)
When xfs_ail_delete() removes an item from the AIL, it calls
xfs_ail_update_finish(), which calls __xfs_ail_assign_tail_lsn():
tail_lsn = __xfs_ail_min_lsn(ailp); // LSN of first item in AIL
if (!tail_lsn)
tail_lsn = ailp->ail_head_lsn; // AIL empty: tail = current head
WRITE_ONCE(log->l_tail_space,
xlog_lsn_sub(log, ailp->ail_head_lsn, tail_lsn));
atomic64_set(&log->l_tail_lsn, tail_lsn);
After updating the tail, xfs_ail_update_finish() calls xfs_log_space_wake(),
which wakes all threads sleeping on l_reserve_head or l_write_head. This
directly unblocks stalled transaction allocations.
The tail can only move forward. It is the minimum LSN of all items still in the AIL. The log space available to new transactions is:
available = l_logsize − (l_tail_space)
= l_logsize − (head_lsn − tail_lsn)
Every item flushed to disk shrinks l_tail_space, freeing space for the next wave
of transactions.
AIL Locking Summary
| Lock | Scope | Notes |
|---|---|---|
ail_lock (spinlock) | All AIL list/state access | Dropped during I/O in inode push |
xfs_buf.b_lock | Individual buffer state | Acquired via trylock only; never spins |
i_pincount / b_pin_count | Pin reference counts | Atomic; checked before attempting push |
The key design rule: ail_lock is never held while waiting for I/O. It is
dropped before xfs_iflush_cluster() and reacquired immediately after, with cursors
protecting traversal safety across the gap.
On-Disk Log Format
Log Record Header (xfs_log_format.h)
struct xlog_rec_header {
__be32 h_magicno; // 0xFEEDbabe
__be32 h_cycle; // wrap count
__be32 h_version; // log version (1 or 2)
__be32 h_len; // data length in bytes
__be64 h_lsn; // this record's LSN
__be64 h_tail_lsn; // oldest uncommitted LSN at write time
__le32 h_crc; // CRC-32c of entire record
__be32 h_num_logops; // count of operations in this record
__be32 h_cycle_data[]; // cycle number embedded in each 512-byte block
uuid_t h_fs_uuid; // filesystem UUID
};
Operation Header
struct xlog_op_header {
__be32 oh_tid; // transaction ID (for grouping ops)
__be32 oh_len; // payload length
__u8 oh_clientid; // XFS_TRANSACTION = 0x69
__u8 oh_flags; // START_TRANS | COMMIT_TRANS | CONTINUE_TRANS
};
Log Item Types
| Type | Value | Description |
|---|---|---|
XFS_LI_INODE | 0x123b | Inode core and data fork |
XFS_LI_BUF | 0x123c | Raw buffer (btree blocks, superblock, etc.) |
XFS_LI_DQUOT | 0x123d | Quota record |
XFS_LI_EFI/EFD | 0x1236/7 | Extent free intent/done |
XFS_LI_RUI/RUD | 0x123a/9 | Rmap update intent/done |
XFS_LI_CUI/CUD | 0x123f/g | Refcount update intent/done |
XFS_LI_BUI/BUD | intent pairs | BMBT update intent/done |
XFS_LI_ATTRI/ATTRD | intent pairs | Xattr update intent/done |
Intent/Done pairs implement a two-phase commit protocol for complex operations that span multiple sub-transactions (e.g., freeing extents requires updating the free space B-tree and the reverse-mapping B-tree). If the filesystem crashes between writing the Intent and the Done record, recovery re-executes the operation from the Intent.
B-tree Splits
A B-tree split is the most log-intensive operation in the XFS metadata path. A single record insertion can trigger a cascade of splits from leaf to root, each allocating a new block and logging multiple buffers. Because reservation sizes are calculated from the worst-case split depth, understanding splits is essential for understanding why XFS log reservations are as large as they are.
When a Split Occurs
XFS B-trees are full B+ trees: every block is kept as full as possible during
insertion. When an insertion targets a block that is already at maximum capacity,
the kernel first tries two cheaper alternatives before resorting to a split
(xfs_btree_make_block_unfull(), xfs_btree.c):
- Left shift (
xfs_btree_lshift()): move the leftmost record to the left sibling if it has space. - Right shift (
xfs_btree_rshift()): move the rightmost record to the right sibling if it has space. - Split (
xfs_btree_split()): only if both siblings are also full.
A split always produces exactly one new block at the current level and returns one new key/pointer pair to the caller, which must then insert that pair into the parent level — potentially triggering another split.
On-Disk Block Format
Every XFS B-tree block on disk begins with struct xfs_btree_block
(libxfs/xfs_btree_format.h):
struct xfs_btree_block {
__be32 bb_magic; // per-btree magic (e.g. XFS_BNOBT_MAGIC)
__be16 bb_level; // 0 = leaf, 1+ = internal node
__be16 bb_numrecs; // number of records/keys currently stored
union {
struct xfs_btree_block_shdr s; // AG-rooted trees (32-bit sibling ptrs)
struct xfs_btree_block_lhdr l; // inode-rooted trees (64-bit sibling ptrs)
} bb_u;
};
Both header variants contain:
| Field | Purpose |
|---|---|
bb_leftsib | Block number of left sibling (or NULLAGBLOCK/NULLFSBLOCK) |
bb_rightsib | Block number of right sibling |
bb_blkno | Physical block address of this block |
bb_lsn | LSN of the last transaction that modified this block |
bb_uuid | Filesystem UUID (guards against cross-filesystem recovery) |
bb_owner | AG number (AG-rooted) or inode number (inode-rooted) |
bb_crc | CRC-32c of the block (recalculated on recovery, not logged) |
Following the header, a block contains either:
- Leaf: a flat array of fixed-size records.
- Internal node: an array of keys followed by an array of
n+1child pointers.
All integer fields are big-endian on disk.
The Split Mechanism: __xfs_btree_split()
The core implementation lives in __xfs_btree_split() (xfs_btree.c). For BMBT
(block map B-tree) splits where no AGF lock is held, a worker-thread wrapper
xfs_btree_split() offloads the call to avoid unbounded kernel stack growth during
recursive allocation; all other tree types call __xfs_btree_split() directly.
Step 1: Allocate the New Right Block
xfs_btree_alloc_block(cur, &lptr, &rptr, stat)
xfs_btree_get_buf_block(cur, &rptr, &right, &rbp)
xfs_btree_init_block_cur(cur, rbp, level, 0)
Block allocation is type-specific:
| B-tree | Source of new block | Side effect logged |
|---|---|---|
| BNOBT / CNTBT | xfs_alloc_get_freelist() — AG free list | AGF header (XFS_AGF_FLFIRST, XFS_AGF_FLCOUNT) |
| INOBT / FINOBT | xfs_alloc_vextent_near_bno() — AG free space | AGF + AGI block counter |
| RMAPBT | xfs_alloc_get_freelist() — AGFL | AGF agf_rmap_blocks, space reservation |
| BMBT | xfs_alloc_vextent_near_bno() — data AG | Inode fork block count |
Every one of these block sources modifies an AG header (AGF or AGI), which is itself
logged as a XFS_LI_BUF item. A split thus always generates at least two logged
buffers before any tree data is touched.
Step 2: Divide Records Between Left and Right
lrecs = xfs_btree_get_numrecs(left)
rrecs = lrecs / 2
if (lrecs is odd && cursor position <= rrecs + 1)
rrecs++ // tilt balance toward right when cursor is nearby
src_index = lrecs - rrecs + 1
xfs_btree_set_numrecs(left, lrecs - rrecs)
xfs_btree_set_numrecs(right, rrecs)
The split point is chosen so that both blocks end up roughly half full. The odd- record tilt biases records toward the block the cursor is about to insert into, minimising the chance of an immediate follow-up split.
For leaf blocks: xfs_btree_copy_recs() copies the upper half of records into
the right block.
For internal nodes: xfs_btree_copy_keys() and xfs_btree_copy_ptrs() copy
the upper half of keys and their associated child pointers.
The key at src_index (the lowest key of the right block) is extracted and returned
to the caller as the split key — the value that must be inserted into the parent
level as the separator between left and right.
Step 3: Log Every Modified Buffer
This is the critical point at which the split becomes durable. The following log operations happen in order:
| What | Fields logged | xfs_btree_log_block() flags |
|---|---|---|
| Right block — all header fields | magic, level, numrecs, both sibling ptrs, blkno, LSN, UUID, owner | XFS_BB_ALL_BITS (excludes bb_crc) |
| Right block — data (leaf) | records 1..rrecs | via xfs_btree_log_recs() |
| Right block — data (node) | keys 1..rrecs, ptrs 1..rrecs | via xfs_btree_log_keys() + xfs_btree_log_ptrs() |
| Left block — changed header | bb_numrecs, bb_rightsib | XFS_BB_NUMRECS | XFS_BB_RIGHTSIB |
| Right-right sibling (if exists) | bb_leftsib | XFS_BB_LEFTSIB |
xfs_btree_log_block() converts the field bitmask to a byte range and calls
xfs_trans_log_buf(), which marks that range dirty in the transaction’s log vector.
xfs_trans_buf_set_type() is called first to stamp the buffer as
XFS_BLFT_BTREE_BUF, which recovery uses to distinguish B-tree blocks from other
buffer types.
The CRC (bb_crc) is deliberately not logged. It is recalculated from the block
contents during recovery using xfs_btree_reada_bufs(), ensuring the stored CRC
always matches what is actually on disk after replay.
Step 4: Update Sibling Chain
Before logging, the sibling doubly-linked list is repaired:
Before split:
[left] ↔ [right-right]
After split:
[left] ↔ [right (new)] ↔ [right-right]
Three pointer writes are needed:
left->bb_rightsib = right(logged as part of left block header)right->bb_leftsib = left(logged as part of right block header,XFS_BB_ALL_BITS)right->bb_rightsib = right-right(logged as part of right block header)right-right->bb_leftsib = right(logged separately:XFS_BB_LEFTSIBonly)
The right-right block read uses xfs_btree_read_buf_block(), which may issue a
synchronous read if the block is not already in the buffer cache. On cold-cache
workloads this is a significant latency source.
Upward Propagation: Recursive Splits
After __xfs_btree_split() returns, the caller (xfs_btree_insrec()) must insert
the split key and right-block pointer into the parent level. If the parent is
also full, it too must split. xfs_btree_insert() drives this loop:
do {
error = xfs_btree_insrec(cur, level, &nptr, &rec, &key, &ncur, &i);
// nptr is non-null if a split occurred at this level
level++;
} while (!xfs_btree_ptr_is_null(cur, &nptr));
Each iteration may allocate one block and log three to five buffers. The loop terminates only when a level has room for the new key/pointer without splitting.
Worst-case depth: on a filesystem with a large allocation group and all optional B-trees enabled, a fully-populated RMAPBT can reach five levels. A single extent allocation that triggers a split at every level logs five new blocks plus five parent block updates plus five AG header updates — thirty or more buffer log items for one allocation.
The per-transaction log reservation must cover this worst case upfront, which is why reservation sizes are computed from tree height and block size at mount time rather than at runtime.
Root Split: Growing the Tree
When the split reaches the root, there is no parent to absorb the new key. The tree must grow one level taller.
AG-Rooted Trees (BNOBT, CNTBT, INOBT, FINOBT, RMAPBT)
xfs_btree_new_root() (xfs_btree.c):
1. Allocate a new block → becomes the new root
2. xfs_btree_set_root(cur, &nptr, +1)
→ update AG header (AGF or AGI) root pointer and level field
→ log AGF/AGI with XFS_AGF_ROOTS | XFS_AGF_LEVELS
3. Initialize new root block (level = old_height, numrecs = 2)
4. Log new root block: XFS_BB_ALL_BITS
5. Copy lowest key of each child into new root keys
6. Log keys: xfs_btree_log_keys(cur, nbp, 1, 2)
7. Write left-child and right-child pointers into new root
8. Log ptrs: xfs_btree_log_ptrs(cur, nbp, 1, 2)
9. Advance cursor: bc_nlevels++
The AG header (AGF or AGI) records the new root block number and the new tree height. On the next mount, XFS reads those fields to reconstruct the cursor starting position without scanning the tree.
Inode-Rooted Trees (BMBT)
xfs_btree_new_iroot() (xfs_btree.c) handles the bmap B-tree, where the root
lives directly inside the inode fork rather than in a separate block:
1. Allocate a new block → receives a copy of current inode-root contents
2. memcpy(new_block, inode_root_data)
Fix bb_blkno in new block to match its physical address
3. Compress the inode fork to hold only the new root (one key + one pointer)
4. Log new child block: XFS_BB_ALL_BITS + records/keys/ptrs
5. Log inode: XFS_ILOG_CORE | xfs_ilog_fbroot(whichfork)
(the inode fork data region is now a single-entry root node)
6. bc_nlevels++
The inode fork has a fixed size defined by its di_forkoff. Once the in-inode root
cannot hold another key/pointer pair even after a split, the inode root gains another
level outward, eventually consuming the entire fork and forcing a fork conversion.
Cursor Tracking Across a Split
The xfs_btree_cur maintains one xfs_btree_level entry per tree level, each
holding a (buffer, position) pair:
struct xfs_btree_level {
struct xfs_buf *bp; // buffer holding the block at this level
uint16_t ptr; // 1-based index of current key/record
};
After __xfs_btree_split() divides the block, the cursor position may have moved
to the right block:
if (cur->bc_levels[level].ptr > lrecs + 1) {
xfs_btree_setbuf(cur, level, rbp); // switch to right block
cur->bc_levels[level].ptr -= lrecs; // adjust position
}
If there are more levels above, a second cursor is duplicated
(xfs_btree_dup_cursor()). One cursor tracks the left child; the duplicated cursor’s
parent-level pointer is incremented by one to track the right child. The insertion
loop in xfs_btree_insert() manages which cursor to use at each level and deletes
the spare when the split chain resolves.
What Gets Logged Per Split Level
Summing the buffer log items produced by one complete split at a single level:
| Buffer | Log items | Condition |
|---|---|---|
| AG header (AGF or AGI) | XFS_LI_BUF | Always — block allocation modifies AG header |
| New right block header | XFS_LI_BUF (XFS_BB_ALL_BITS) | Always |
| New right block data (recs/keys/ptrs) | XFS_LI_BUF | Always |
| Left block header | XFS_LI_BUF (XFS_BB_NUMRECS | XFS_BB_RIGHTSIB) | Always |
| Right-right sibling header | XFS_LI_BUF (XFS_BB_LEFTSIB) | Only if right-right exists |
That is four or five XFS_LI_BUF items per split level, plus the AG header
items from block allocation, which may themselves modify additional blocks (e.g.,
the AGFL block list used by BNOBT/RMAPBT to store free blocks). In a five-level tree
where every level splits, a single insertion can log twenty or more distinct buffers
before the transaction commits.
Crash Recovery of a Split
Because every buffer modified during a split is logged before the transaction commits, crash recovery is straightforward:
- Crash before commit: no log records for this transaction are durable. The pre-split blocks are unmodified. The newly allocated block may appear in the block allocation structures but will be reclaimed by the space recovery pass.
- Crash after commit: all log records are durable. Recovery replays each buffer item in order, restoring every block to its post-split state. The B-tree is structurally consistent at the end of replay.
There are no intent records for splits. A split is not a deferred operation: it is atomic within the transaction that triggered it. Either all split records are committed together, or none of them are.
Crash Recovery
xfs_log_recover.c implements recovery in two passes.
Pass 1: Log Scanning
xlog_recover() walks the log from tail_lsn forward:
- Reads each log record header.
- Validates magic number and CRC.
- Groups operation headers by transaction ID.
- Builds an in-memory map of all transactions present in the log.
Pass 2: Replay
xlog_recover_commit_trans() replays each complete transaction:
- For each log item, call
xlog_recover_commit_buffer/inode/dquot(). - Overwrite on-disk metadata with the logged versions.
- For intent items (EFI, RUI, etc.), reconstruct the pending operation and schedule it for Phase 2 completion via deferred operations.
xlog_recover_finish() processes all deferred operations, completing any
partially-done multi-step operations (extent frees, rmap updates, etc.).
Invariant enforced by design: a CIL checkpoint must be smaller than half the total log size. This guarantees that at least one full checkpoint is always present in the log, making partial-write crashes safe.
Locking Hierarchy
Violating this order causes deadlock.
1. xfs_mount.m_sb_lock (filesystem-wide, rarely held)
2. xfs_buf.b_lock (individual buffer locks)
3. xfs_inode.i_lock (inode lock)
4. xlog.l_icloglock (spinlock, iclog state machine)
5. xfs_cil.xc_ctx_lock (rwsem, CIL context switch)
6. xfs_cil.xc_push_lock (spinlock, checkpoint ordering list)
7. xfs_ail.ail_lock (spinlock, AIL list)
xc_ctx_lock is a sleeping rwsem, not a spinlock, specifically to avoid holding a
spinlock during log I/O submission, which can sleep.
Performance Characteristics and Bottlenecks
Batching Efficiency (CIL)
The CIL’s primary value is write amplification reduction. Without delayed
logging, each fsync or log force flushes all dirty items individually. With CIL,
items modified 100 times between two checkpoints are written to disk once, at
their final state. In workloads with heavy relogging (directory updates, quota
updates), this can reduce log I/O by an order of magnitude.
Bottleneck: If the CIL is too small (< 8 MB on a busy filesystem), background pushes fire too frequently, destroying the batching benefit and driving up log I/O.
Per-CPU CIL Aggregation
Transaction commits add items to per-CPU pending lists (xc_pcp) to eliminate
contention on a single lock. Space accounting uses per-CPU counters until the soft
limit is approached, at which point it transitions to atomic operations and wakes the
push worker.
Bottleneck: On workloads with very many small transactions (e.g., millions of small file creates), the per-CPU-to-atomic transition point creates a serialization spike. Threads pile up in
xlog_cil_commit()contending onxc_ctx_lockwrite acquisition during the context switch.
Grant Head Waiters
When the log is full (write head nearly meets the tail), new transactions block in
xlog_grant_head_wait() on a FIFO wait queue.
Bottleneck — log tail pinning: The tail can only advance when the AIL empties items. The AIL can only empty items when their buffers are written to disk. If the storage device is slow, the log fills up and all new transactions stall. This is the primary throughput bottleneck on write-heavy workloads on slow devices.
Bottleneck — reservation overestimation: Reservations are computed for the worst case (maximum B-tree depth). On a mostly-empty filesystem, actual usage is much less, but the reservation holds the full amount until released. This reduces parallelism on small logs.
Iclog Contention (l_icloglock)
Every thread completing a CIL commit must briefly hold l_icloglock to copy its log
vectors into the current iclog and advance the write cursor. On many-core systems
(32+ CPUs), this spinlock becomes a serialization point under high log bandwidth.
Bottleneck: Large CIL checkpoints writing megabytes of log data while holding
l_icloglockfor each 32 KB iclog block starve concurrent threads trying to start new transactions.
AIL Push Rate and Tail Stall
xfsaild is a single-threaded daemon. On systems with many concurrent metadata
writers, it must push items fast enough to keep the tail advancing ahead of the
write head.
Bottleneck — device throughput: If the block device cannot sustain the required writeback rate,
xfsaildbuilds up a backlog, the AIL grows, the tail stalls, the write head catches the tail, and transaction allocation blocks — a global freeze. The only remedy is faster storage, a larger log, or reducing metadata write amplification.
Bottleneck — pinned items: Items held by long-running or stalled transactions cannot be pushed regardless of device speed. When
xfsaildencounters more than 100 pinned items in a row it backs off (20–50 ms sleep). If the items remain pinned across many rounds,ail_log_flushaccumulates and each new push round opens with a forced CIL flush to try to unpin them. A transaction that holds its locks too long effectively pins the log tail and starves all other writers.
Bottleneck — buffer lock contention:
xfsaildusestrylockon all buffers and returnsXFS_ITEM_LOCKEDimmediately if the lock is unavailable. Under heavy concurrent writeback, many buffers may be locked by page writeback or other kernel paths. Items returningLOCKEDcount against thestuckthreshold (100 items), triggering the backoff before the target LSN is reached.
Bottleneck — single-threaded design:
xfsaildprocesses the AIL serially. Each iteration submits buffers viaxfs_buf_delwri_submit_nowait(), which is asynchronous, but the traversal itself is sequential. On workloads that produce millions of small dirty metadata items, the daemon can spend more time traversing the list than the device spends doing I/O. There is no parallelism within a single push round.
Bottleneck — cluster flush overhead: Inode pushes call
xfs_iflush_cluster()which formats all inodes in a buffer cluster. While this amortizes I/O, it also meansxfsailddrops and reacquiresail_lockfor every inode buffer, and any cursor invalidation during that window forces a restart of the traversal from the AIL minimum. On a filesystem with millions of recently-modified inodes spread across many clusters, this restart overhead can significantly slow the effective push rate.
Observable symptoms of a stalled tail:
xfs_log_forcelatency increases (callers sleeping onl_write_head).xfs_buf_delwri_submit_nowaitreturns non-zero repeatedly (setsail_log_flusheach time), causing redundant CIL flushes./proc/fs/xfs/statcountersxs_push_ail_pinnedandxs_push_ail_lockedgrow faster thanxs_push_ail_success.xfsaildwakes withtout=20continuously (>90% contention threshold crossed).
Tuning levers:
- Larger log: more space between head and tail gives
xfsaildmore time before the write head catches the tail. - Dedicated log device (separate fast NVMe): isolates log writes from data writeback, reducing contention on the device queue.
vm.dirty_ratio/vm.dirty_background_ratio: reducing the dirty page ratio limits how many buffers can be in-flight at once, reducing lock contention seen byxfsaild.
Checkpoint Ordering Serialization
xlog_cil_order_write() (xfs_log_cil.c) ensures commit records are written in
sequence order. When two concurrent checkpoints race, the higher-sequence one must
wait for the lower-sequence one to establish its commit_lsn before writing its own
commit record.
Bottleneck: Under extreme concurrency with many small checkpoints firing in rapid succession, checkpoint ordering serialization limits the rate at which new commit LSNs can be established, capping throughput in the log-write path.
Recovery Time
Recovery time is proportional to the amount of data between tail_lsn and
head_lsn at the time of crash. A larger log retains more history, meaning more
data to replay. On systems with very large logs (hundreds of GB) and high write rates
before the crash, recovery can take minutes.
Bottlenecks and Write Amplification
Write amplification in XFS logging occurs at several independent layers. Each layer multiplies the number of actual device writes relative to the application-level operation that triggered them. Understanding which layer is responsible for observed I/O load is essential for diagnosis.
Write Amplification Taxonomy
Layer 1: Fundamental WAL Amplification
Every metadata modification is written twice: once sequentially to the log, and
once in-place to the metadata location on disk. This is the irreducible cost of
crash consistency via WAL. A single mkdir that modifies an inode, a directory
block, and two AGF entries produces at minimum four log writes and four eventual
on-disk writes — eight device writes for four logical changes.
Application write
└─ metadata change
├─ → log write (sequential, via iclog)
└─ → on-disk write (random, via AIL writeback)
The log write is sequential and cheap per-byte. The on-disk write is random and expensive per-operation. For metadata-heavy workloads on rotational storage the random on-disk writes dominate; on NVMe the log bandwidth is more often the limit.
Layer 2: Relogging Amplification (Pre-CIL)
Before delayed logging, every transaction that modified an already-logged item wrote the item to the log again in full, even if the change was a single byte. A hot inode touched by 1 000 transactions before being flushed to disk would appear 1 000 times in the log. Log space consumption was proportional to transaction count, not to the number of distinct objects.
CIL eliminates this at the log level: the item is formatted once per checkpoint regardless of how many transactions modified it within that checkpoint window. The reduction in log write amplification depends entirely on the relogging rate. On workloads with heavy relogging (directory entry updates, quota tracking, allocation group headers), CIL can reduce log write volume by one to two orders of magnitude.
Layer 3: Shadow Buffer Copy Amplification
CIL introduces one additional in-memory copy per commit. Each item is formatted from
its live in-memory representation into a shadow buffer (xfs_log_vec.lv_buf)
before being added to the CIL. This decouples the item from the log write so the
item can be unlocked immediately, but it means every committed item is represented in
memory at least twice: once as the live object (inode, buffer) and once as the
formatted shadow. The shadow is later copied into the iclog when the CIL pushes.
Memory path for a single logged inode:
xfs_inode (in memory)
→ iop_format() → lv_buf (shadow buffer, CIL holds it)
→ xlog_write() → iclog data buffer (ring)
→ disk
Three copies before the data reaches the log device. This amplification is intentional: it removes the need to hold any lock on the live object during log I/O, enabling the parallelism that makes delayed logging viable.
Layer 4: Metadata Cascade Amplification (B-tree Fan-out)
A single application-visible operation triggers a cascade of internal metadata changes, each of which must be logged independently. The worst case occurs during extent allocation on a filesystem with all optional B-trees enabled:
| Operation step | Items logged |
|---|---|
| Inode size/extent count update | XFS_LI_INODE |
| BMBT (extent map B-tree) block | XFS_LI_BUF × (tree height) |
| AGF header update | XFS_LI_BUF |
| Free space B-tree by block (BNOBT) | XFS_LI_BUF × (split depth) |
| Free space B-tree by size (CNTBT) | XFS_LI_BUF × (split depth) |
| Rmap B-tree (RMAPBT, if enabled) | XFS_LI_BUF × (split depth) |
| Refcount B-tree (REFCBT, if enabled) | XFS_LI_BUF × (split depth) |
| AGI header (if inode allocation) | XFS_LI_BUF |
| Inode B-tree (INOBT/FINOBT) | XFS_LI_BUF × (split depth) |
A single fallocate call on a filesystem with rmapbt and refcountbt enabled can
log 20–40 buffer items. Each B-tree split creates a new block that must also be
logged. The reservation system accounts for this worst case, which is why per-
transaction reservations are large relative to the actual bytes changed.
Layer 5: Intent/Done Record Overhead
Multi-step operations (extent free, rmap update, refcount update, attribute write) write a pair of log records — an Intent before the operation and a Done after. This ensures recovery can detect and complete partial operations. The overhead is two additional log records per complex sub-operation:
EFI (Extent Free Intent) ← written before freeing extent
→ btree updates (AGF, BNOBT, CNTBT, RMAPBT, REFCBT)
EFD (Extent Free Done) ← written after btree updates complete
On workloads that perform many small file deletions (e.g. log rotation, build artifact cleanup), Intent/Done pairs can account for a significant fraction of log traffic. A delete of a 100-extent file generates 100 EFI/EFD pairs plus all associated B-tree buffer logs.
Layer 6: Iclog Block Padding
Log records are written in units of 512-byte blocks and padded to the next block boundary. Small transactions that log only a few hundred bytes waste the remainder of the block. On workloads with many small transactions the padding overhead can reach 30–50% of raw log bandwidth, effectively shrinking the usable log size.
The CIL largely mitigates this by batching many small transactions into a single large checkpoint record. Padding waste is then amortized across the checkpoint rather than per-transaction.
Bottleneck Catalog
Each bottleneck is described with its root cause, how it manifests in observable metrics, and what can be done to mitigate it.
B1: Log Full — Grant Head Stall
Root cause: The write head has caught up to the tail. No physical log space
remains for new transactions. All calls to xlog_grant_head_check() block on the
l_write_head FIFO wait queue.
Cause chain:
Device too slow → AIL drain lags → tail does not advance
→ write head catches tail → xlog_grant_head_wait() blocks all writers
Symptoms:
- All application threads stall in
xfs_log_reserve()simultaneously — a global filesystem freeze from the application’s perspective. dmesgmay showXFS: xlog_grant_log_space: sleepif debug logging enabled.iostatshows log device at 100% utilisation with very low metadata device I/O (metadata writes are blocked waiting for log space)./proc/fs/xfs/stat:xs_trans_ailstalled;xs_push_ail_successnear zero.
Mitigations:
- Increase log size (
mkfs.xfs -l size=...or external log device). - Move the log to a dedicated faster device (NVMe vs. HDD).
- Reduce B-tree fan-out amplification by enabling
bigtime,nrext64, or choosing a larger block size to pack more records per B-tree node. - Reduce the number of enabled optional B-trees if rmap/reflink are not required.
B2: CIL Context Switch Contention
Root cause: The CIL push worker acquires xc_ctx_lock as a writer to swap
the live context. During this window, all concurrent xlog_cil_commit() calls block
waiting for the read lock. On many-core systems with high transaction rates the
context switch becomes a serialisation barrier.
Symptoms:
- CPU profiles show many threads spinning or sleeping in
xlog_cil_commit(). - Short bursts of very high lock wait time correlate with checkpoint boundaries.
- Transaction commit latency has a periodic spike pattern matching the CIL push interval (every few hundred milliseconds under load).
Mitigations:
- The CIL is already tuned to minimise the write-lock hold time (context is swapped
then lock released immediately). The main lever is reducing push frequency by
ensuring the log is large enough that the CIL soft limit (
XLOG_CIL_SPACE_LIMIT, ~12.5% of log) is not hit too often. - Workloads that issue many synchronous
fsynccalls force CIL pushes on every call. Batchingfsync(e.g. usingsync_file_rangeor application-level buffering) reduces push frequency.
B3: Iclog Spinlock Serialisation (l_icloglock)
Root cause: Every thread writing to an iclog must hold l_icloglock for the
duration of the copy. The lock is a raw spinlock. On systems with 32+ CPUs all
running concurrent CIL push workers, the spinlock degrades to a bottleneck.
Symptoms:
perforftraceshows high time in_raw_spin_lockcalled fromxlog_write_iclog()orxlog_state_get_iclog_space().- Log write bandwidth plateaus below the device’s sequential write capacity.
- Adding more CPUs does not improve log throughput.
Mitigations:
- Use larger iclogs (
mkfs.xfs -l version=2,size=...,su=262144sets 256 KB iclogs). Larger iclogs mean fewer lock acquisitions per unit of log data. - Reduce the number of concurrent CIL pushes by ensuring workload transactions are large enough to batch well before hitting the CIL limit.
B4: Reservation Overestimation on Small Logs
Root cause: Each transaction type holds a worst-case reservation for the entire duration of the transaction, even if the actual log usage is a fraction of that. On a small log (< 256 MB), the sum of all in-flight reservations can exhaust the reserve head even when the physical log has space, causing false stalls.
Symptoms:
xlog_grant_head_check()blocks onl_reserve_headeven thoughl_write_headhas available space.- Log device utilisation is low but transaction latency is high.
- Reducing concurrent writer count relieves the stall.
Mitigations:
- Increase log size. Reservations are a fixed fraction of log size; a larger log accommodates more concurrent in-flight transactions.
- Avoid small logs on high-concurrency filesystems. The minimum practical log size for a busy filesystem is typically 512 MB; 1–2 GB is common on production systems.
B5: AIL Tail Stall — Pinned Items
Root cause: Items in the AIL that are still pinned by in-flight CIL transactions
cannot be flushed. If the CIL checkpoint does not complete quickly enough (e.g.,
because log I/O is slow), the tail cannot advance. xfsaild counts pinned items
against the stuck threshold and backs off after 100 consecutive pinned items.
Symptoms:
/proc/fs/xfs/stat:xs_push_ail_pinneddominates overxs_push_ail_success.xfsaildsleep time is consistently 20–50 ms (backoff mode).ail_log_flushcounter increments rapidly, causing redundant CIL flushes.- Log I/O latency is high (slow log device or iclog congestion).
Mitigations:
- Faster log device reduces the time between CIL commit and iclog I/O completion, unpinning items sooner.
- If the workload uses explicit
fsync, confirm that it is not being called at a rate that prevents the CIL from batching effectively.
B6: AIL Tail Stall — Buffer Lock Contention
Root cause: xfsaild uses trylock on all buffers. Under heavy concurrent
writeback from the page cache or other kernel paths, many metadata buffers are
already locked when xfsaild tries to acquire them. Each failure increments the
stuck counter toward the 100-item backoff threshold.
Symptoms:
/proc/fs/xfs/stat:xs_push_ail_lockedgrows alongsidexs_push_ail_pinned.- High
iowaiton the metadata device during writeback storms. xfsaildalternates between 0 ms (making progress) and 20 ms (backoff) with no clear pattern.
Mitigations:
- Reduce concurrent writeback pressure via
vm.dirty_background_ratioandvm.dirty_ratio. - On NVMe, increase the
nr_requestsqueue depth to absorb more concurrent I/Os without serialising at the block layer.
B7: AIL Single-Threaded Traversal
Root cause: xfsaild is one thread per filesystem. The AIL traversal loop is
sequential. On workloads that accumulate millions of dirty metadata items (e.g. large
rsync, git clone of a large repository, database checkpoint), the traversal itself
consumes significant CPU time before items are submitted.
Symptoms:
xfsaildCPU usage is consistently high (near 100% of one core).- Log device I/O queue is not saturated —
xfsaildis the bottleneck, not the device. - Cursor restarts are frequent:
xfsaildrepeatedly restarts from the AIL minimum due to concurrent deletions during inode cluster flushes.
Mitigations:
- There is no kernel-level tuning knob to parallelise
xfsaild. The mitigation is to reduce the number of items in the AIL at any one time by ensuring the CIL pushes frequently and items are promptly written to disk. - Increasing
vm.dirty_expire_centisecsdelays the page cache writeback that competes withxfsaild, reducing cursor invalidation interference.
B8: Checkpoint Ordering Stall
Root cause: xlog_cil_order_write() enforces that commit records appear in
strictly ascending checkpoint sequence order. A slow checkpoint (due to large
checkpoint size or log I/O contention) blocks all higher-sequence checkpoints from
writing their commit records, even if their data has already been written to iclogs.
Symptoms:
- Multiple CIL push workers are stalled in
xlog_cil_order_write()waiting onxc_commit_wait. - Log device appears idle despite pending checkpoint data.
- Checkpoint commit latency increases proportionally to checkpoint I/O latency.
Mitigations:
- Reduce checkpoint size by reducing the CIL soft limit (not directly tunable at runtime; requires log size adjustment since the limit is a fraction of log size).
- Faster log device reduces per-checkpoint I/O time, shortening the ordering wait.
B9: Write Amplification from Optional B-Trees
Root cause: Enabling rmapbt (reverse mapping) and reflink (reference
counting) adds two additional B-trees that must be updated on every extent
allocation, deallocation, and CoW operation. Each B-tree update is a separate logged
buffer item. On workloads with high extent churn, these trees double or triple the
number of buffer items logged per operation.
Symptoms:
- Log write bandwidth is significantly higher after enabling reflink or rmapbt compared to a plain filesystem.
- Per-transaction reservation sizes are larger (visible via
xfs_logprint). - Extent allocation operations are slower under concurrency due to higher per- transaction lock hold times.
Mitigations:
- Do not enable
rmapbtorreflinkif the workload does not require them. These features cannot be disabled aftermkfs. - Use a larger block size to increase B-tree node fanout, reducing tree height and therefore split frequency.
B10: Recovery Time from Large Logs
Root cause: Recovery replays every log record between tail_lsn and head_lsn
at crash time. A larger log retains more history. A filesystem that was writing
heavily immediately before the crash will have a full or nearly-full log to replay.
Symptoms:
- Mount time is minutes rather than seconds after an unclean shutdown.
- Recovery I/O is visible on the log device during mount.
dmesgshowsXFS: starting recoveryfollowed by a long gap beforeXFS: Ending recovery.
Mitigations:
- Use
barrier=1(default) to ensure log records are committed before the device acknowledges the write, keeping the recovery window bounded. - A dedicated log device with lower write latency reduces the time to write
checkpoints, keeping
tail_lsncloser tohead_lsnat any given moment (less to replay). - Do not artificially inflate log size beyond what is needed. A log larger than necessary does not improve steady-state performance and increases worst-case recovery time.
Amplification Summary
| Layer | What is amplified | CIL mitigation |
|---|---|---|
| Fundamental WAL | Every metadata write appears twice (log + disk) | None — inherent to WAL |
| Relogging | Hot items logged once per transaction | Yes — once per checkpoint |
| Shadow buffer copies | 3 in-memory copies before disk | Unavoidable cost of lock-free commit |
| B-tree cascade | 10–40 buffer items per file operation | Partial — items are batched per checkpoint |
| Intent/Done pairs | 2 log records per multi-step operation | Partial — both records batched in checkpoint |
| Iclog padding | Up to 512 bytes wasted per transaction | Yes — padding amortised across checkpoint |
| Optional B-trees | 2× log traffic with rmapbt + reflink | None — structural overhead |
Summary of Critical Paths
| Path | Key bottleneck |
|---|---|
| Transaction allocation | l_reserve_head FIFO wait when log is full |
| CIL commit | xc_ctx_lock contention during context switch |
| CIL push → iclog write | l_icloglock on every 32 KB iclog block |
| Iclog I/O completion | Block device latency |
| AIL push — device throughput | xfsaild delwri queue depth vs. device bandwidth |
| AIL push — pinned items | Long-running transactions pin tail; ail_log_flush triggers CIL force |
| AIL push — buffer contention | trylock failures accumulate; 100-item stuck threshold triggers backoff |
| AIL push — single-threaded | Sequential traversal; cursor restarts on concurrent deletion |
| AIL cluster flush | ail_lock drop/reacquire per inode cluster; cursor restart on remove |
| Log tail advance | __xfs_ail_assign_tail_lsn() wakes grant head waiters; stalls if AIL never drains |
| Recovery | Log size × write rate at crash time |
Understanding these seven paths and their limiting factors is the foundation for diagnosing and resolving XFS performance problems on write-intensive workloads.
Source References
| File | Purpose |
|---|---|
fs/xfs/xfs_log.c | Main log manager, iclog state machine |
fs/xfs/xfs_log_cil.c | Delayed logging, CIL push worker |
fs/xfs/xfs_trans_ail.c | AIL daemon (xfsaild), push loop, tail assignment, cursor management |
fs/xfs/xfs_trans_priv.h | xfs_ail and xfs_ail_cursor structure definitions |
fs/xfs/xfs_inode_item.c | xfs_inode_item_push(), cluster flush dispatch |
fs/xfs/xfs_buf_item.c | xfs_buf_item_push(), buffer trylock and delwri queue |
fs/xfs/xfs_log_recover.c | Crash recovery, two-pass replay |
fs/xfs/xfs_log_priv.h | Internal structures (xlog, xlog_in_core, xfs_cil) |
fs/xfs/xfs_log.h | Public logging API |
fs/xfs/libxfs/xfs_log_format.h | On-disk format (xlog_rec_header, item types) |
Documentation/filesystems/xfs/xfs-delayed-logging-design.rst | Authoritative design document |
XFS Dirty Log: head/tail error analysis
Context
When _check_xfs_filesystem runs after a test that triggers ENOSPC during
mmap copy-on-write (e.g. generic/173), it can report two errors:
_check_xfs_filesystem: filesystem on /dev/loop1 has dirty log
_check_xfs_filesystem: filesystem on /dev/loop1 is inconsistent (r)
What the errors mean
Dirty log
_check_xfs_filesystem unmounts the filesystem and then runs xfs_logprint -t.
A clean unmount should flush all dirty buffers, drain the AIL, and write a clean
log marker. If the log state is still <DIRTY> after unmount, it means XFS was
force-shut down before the unmount could checkpoint the log — preventing it from
writing the clean marker.
Inconsistent (r)
xfs_repair -n (read-only mode) reads the raw on-disk data structures. Because
the AIL never flushed the logged changes to those structures, the on-disk state
is the pre-transaction state — inconsistent with what the log says should be
there. xfs_repair -n cannot replay the log to reconcile them, so it flags the
filesystem as inconsistent.
A plain xfs_repair (without -n) would replay the log first and would very
likely find no actual corruption.
Example log output
xfs_logprint:
data device: 0x701
log device: 0x701 daddr: 327728 length: 131072
log tail: 40 head: 46 state: <DIRTY>
- Internal log — log and data are on the same device (
0x701). tail: 40, head: 46— 6 blocks of live log records sit between tail and head. These records were written to the on-disk log but the corresponding buffer writes (AGF, btree blocks, inodes) never reached the data structures. On the next mount XFS would replay them.
Transactions in the log
Transaction 1 — LSN (cycle 1, block 40), tid 0x31cd442c, 11 items
| Item | Detail |
|---|---|
AGF buffer at blkno 0x1 | AG 0 free space manager — space was allocated or freed |
3 data buffers at 0x8, 0x10, 0x28 | Free space or inode btree blocks modified by the allocation |
Inode 0x84 (132), flags 0x5, dsize 48 | Core + data fork extents — a file whose extent map was being updated |
Consistent with a CoW block allocation in progress: AGF modified, btree blocks updated, and the target inode’s extent map changed.
Transaction 2 — LSN (cycle 1, block 44), tid 0xa2d07cca, 2 items
| Item | Detail |
|---|---|
Inode 0x83 (131), flags 0x1, dsize 0 | Core only, no extent data — an inode whose data fork was cleared or reset |
Inode log item anatomy
Using inode 0x84 from transaction 1 as an example:
INO: cnt:3 total:3 a:0xaaaaef94fe40 len:56 a:0xaaaaef94fef0 len:176 a:0xaaaaef94ffb0 len:48
INODE: #regs:3 ino:0x84 flags:0x5 dsize:48
CORE inode:
DATA FORK EXTENTS inode data:
The inode log item is composed of 3 memory regions copied into the log:
| Address | Length | Content |
|---|---|---|
0xaaaaef94fe40 | 56 bytes | xfs_inode_log_format — header describing which parts of the inode were dirtied |
0xaaaaef94fef0 | 176 bytes | Inode core (xfs_dinode) — timestamps, size, link count, etc. |
0xaaaaef94ffb0 | 48 bytes | Data fork extent records |
flags:0x5 = XFS_ILOG_CORE (0x1) | XFS_ILOG_DEXT (0x4) — both the inode
core and the data fork extent list were dirtied by this transaction.
dsize:48 — 48 bytes of extent data = 3 extents (each xfs_bmbt_rec_t
is 16 bytes). These are the packed extent records describing the new block
mappings for the CoW operation.
Summary of what happened
- The test fills the filesystem and then attempts an mmap CoW write with no free space.
- XFS starts a transaction: allocates CoW blocks (modifying the AGF and btree blocks) and updates the target inode’s extent map.
- ENOSPC is hit mid-flight. XFS force-shuts down the filesystem to avoid leaving metadata in a half-updated state.
- The transaction records are committed to the on-disk log (that is why
xfs_logprintcan see them), but the AIL never flushes the modified buffers to their actual on-disk locations — the log tail is never pushed. - Force-shutdown prevents any further log writes, so the clean log marker cannot be written on unmount.
xfs_logprintsees<DIRTY>andxfs_repair -nsees stale on-disk structures that do not match the logged intent.
NFS
Filesystem
NFS with KRB5
Setup KDC:
- Install the required packages for the KDC.
[root@kdc-server ~]# dnf install krb5-libs krb5-server krb5-workstation
- Edit the
/etc/krb5.conf.
# To opt out of the system crypto-policies configuration of krb5, remove the
# symlink at /etc/krb5.conf.d/crypto-policies which will not be recreated.
includedir /etc/krb5.conf.d/
[logging]
default = FILE:/var/log/krb5libs.log
kdc = FILE:/var/log/krb5kdc.log
admin_server = FILE:/var/log/kadmind.log
[libdefaults]
dns_lookup_realm = false
ticket_lifetime = 24h
renew_lifetime = 7d
forwardable = true
rdns = false
pkinit_anchors = FILE:/etc/pki/tls/certs/ca-bundle.crt
spake_preauth_groups = edwards25519
dns_canonicalize_hostname = fallback
qualify_shortname = ""
default_realm = EXAMPLE.COM
default_ccache_name = KEYRING:persistent:%{uid}
[realms]
EXAMPLE.COM = {
kdc = kdc.example.com
admin_server = kdc.example.com
}
[domain_realm]
.example.com = EXAMPLE.COM
example.com = EXAMPLE.COM
- Create the database using the kdb5_util.
[root@kdc-server ~]# kdb5_util create -s
- Set ACL in the
/var/kerberos/krb5kdc/kadm5.acl. Bellow settings allows anyone with secodary admin principal to have full administrative access for example:user/admin@EXAMPLE.COM`
*/admin@EXAMPLE.COM *
- Create the first principal using kadmin.local at the KDC terminal:
[root@kdc-server ~]# kadmin.local -q "addprinc user/admin"
- Satrt krb5kdc and kadmin.
[root@kdc-server ~]# systemctl enable --now kadmin.service krb5kdc
NFS server configuration.
- Install packages for kerberos client and NFS server
[root@nfs-server ~]# dnf install krb5-workstation nfs-utils
- NFSv4 idmapping becomes much more important to have with Kerberos. Both the
server and the clients should have the same idmapping domain configured. In
the
/etc/idmapd.confset the domain to your kerberos realm.
[General]
Domain = example.com
- Each NFS server needs a Kerberos principal for
nfs/server.fqdnto be created on the KDC, and its keys added to the server’s /etc/krb5.keytab.
[root@nfs-server ~]# kadmin -p username/admin
Password for username/admin@EXAMPLE.COM: ***********
kadmin: addprinc -nokey nfs/nfs-server.example.com
kadmin: addprinc -nokey host/nfs-server.example.com
kadmin: ktadd nfs/nfs-server.example.com
kadmin: ktadd host/nfs-server.example.com
[root@nfs-server ~]# klist -ke
[root@nfs-server ~]# klist -ke
Keytab name: FILE:/etc/krb5.keytab
KVNO Principal
---- -------------------------------------------------------------------
1 host/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha384-192
1 host/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha256-128)
1 host/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha1-96)
1 host/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha1-96)
1 host/nfs-server.example.com@EXAMPLE.COM (camellia256-cts-cmac)
1 host/nfs-server.example.com@EXAMPLE.COM (camellia128-cts-cmac)
1 host/nfs-server.example.com@EXAMPLE.COM (DEPRECATED:arcfour-hmac)
1 nfs/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha384-192)
1 nfs/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha256-128)
1 nfs/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha1-96)
1 nfs/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha1-96)
1 nfs/nfs-server.example.com@EXAMPLE.COM (camellia256-cts-cmac)
1 nfs/nfs-server.example.com@EXAMPLE.COM (camellia128-cts-cmac)
1 nfs/nfs-server.example.com@EXAMPLE.COM (DEPRECATED:arcfour-hmac)
- Enable and start the gssproxy.service
[root@nfs-server ~]# systemctl enable --now gssproxy.service
NFS client configuration
-
Install the
nfs-utilsandkrb5-workstationas on the NFS server and create same configuration filekrb5.conf. -
Add nfs-client to the kerberos.
[root@nfs-server ~]# kadmin -p username/admin
Password for username/admin@EXAMPLE.COM: ***********
kadmin: addprinc -nokey host/nfs-client.example.com
kadmin: ktadd host/nfs-client.example.com
[root@nfs-server ~]# klist -ke
Benchmarks
Benchmark of NFSv4.2 with different security context.
Environment
NFS Server and KDC:
- OS: Fedora 40 Qemu KVM virtual machine.
- 20 CPUs Intel® Xeon® Gold 5215 CPU @ 2.50GHz.
- 250GB memory 100GB used for
/dev/pmem0emulation. - CPUs pinned on NUMA node 0 (NFS clients are pinned to NUMA node 1).
- All configs are left in default values including nfs.conf.
- NFS export is placed on XFS backing up
/dev/pmem0and not using DAX option. - Virtio isolated network.
NFS clients
- OS: version from RHEL 7 to RHEL 9 Qemu KVM
- 8 CPUs Intel® Xeon® Gold 5215 CPU @ 2.50GHz with 8GB of RAM.
- Virtio isolated network.
- Default NFS mount options:
fedora-kdc.example.com:/mnt/nfs /mnt/sys nfs4 rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1 0 0
fedora-kdc.example.com:/mnt/nfs /mnt/krb5 nfs4 rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1 0 0
fedora-kdc.example.com:/mnt/nfs /mnt/krb5i nfs4 rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5i,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1 0 0
fedora-kdc.example.com:/mnt/nfs /mnt/krb5p nfs4 rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5p,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1 0 0
Base line testing with follwing fio configuration:
# cat nfs_08k_rand_write.ini
[global]
direct=1
ioengine=libaio
bs=8k
gtod_reduce=1
size=2G
iodepth=128
rw=randwrite
group_reporting=1
numjobs=4
filename_format=/mnt/$jobname/nfs.$jobnum
[sys]
stonewall=1
[krb5]
stonewall=1
[krb5i]
stonewall=1
[krb5p]
stonewall=1
Results on the NFS server writing directly to the local filesystem. The stats on
diskstats are zero as /dev/pmem0 does not report anything to
/proc/diskstats.
[root@fedora-kdc ~]# fio xfs_08k_rand_write.ini
nfs: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
fio-3.36
Starting 4 processes
Jobs: 4 (f=4): [w(4)][94.1%][w=7443MiB/s][w=953k IOPS][eta 00m:02s]
nfs: (groupid=0, jobs=4): err= 0: pid=8415: Thu Aug 15 11:22:14 2024
write: IOPS=331k, BW=2586MiB/s (2712MB/s)(80.0GiB/31677msec); 0 zone resets
bw ( MiB/s): min= 179, max= 7466, per=99.74%, avg=2579.39, stdev=829.78, samples=252
iops : min=22962, max=955700, avg=330161.65, stdev=106212.09, samples=252
cpu : usr=6.09%, sys=48.51%, ctx=695334, majf=0, minf=31
IO depths : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
submit : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
complete : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
issued rwts: total=0,10485760,0,0 short=0,0,0,0 dropped=0,0,0,0
latency : target=0, window=0, percentile=100.00%, depth=128
Run status group 0 (all jobs):
WRITE: bw=2586MiB/s (2712MB/s), 2586MiB/s-2586MiB/s (2712MB/s-2712MB/s), io=80.0GiB (85.9GB), run=31677-31677msec
Disk stats (read/write):
pmem0: ios=0/0, sectors=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
RHEL 7:
- Encryption type: aes256-cts-hmac-sha1-96
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 0 (all jobs):
WRITE: bw=179MiB/s (187MB/s), 179MiB/s-179MiB/s (187MB/s-187MB/s), io=8192MiB (8590MB), run=45878-45878msec
Run status group 1 (all jobs):
WRITE: bw=173MiB/s (182MB/s), 173MiB/s-173MiB/s (182MB/s-182MB/s), io=8192MiB (8590MB), run=47316-47316msec
Run status group 2 (all jobs):
WRITE: bw=130MiB/s (137MB/s), 130MiB/s-130MiB/s (137MB/s-137MB/s), io=8192MiB (8590MB), run=62825-62825msec
Run status group 3 (all jobs):
WRITE: bw=81.9MiB/s (85.8MB/s), 81.9MiB/s-81.9MiB/s (85.8MB/s-85.8MB/s), io=8192MiB (8590MB), run=100064-100064msec
RHEL 8:
- Encryption type: aes256-cts-hmac-sha384-192
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 1 (all jobs):
WRITE: bw=176MiB/s (185MB/s), 176MiB/s-176MiB/s (185MB/s-185MB/s), io=8192MiB (8590MB), run=46488-46488msec
Run status group 2 (all jobs):
WRITE: bw=154MiB/s (161MB/s), 154MiB/s-154MiB/s (161MB/s-161MB/s), io=8192MiB (8590MB), run=53239-53239msec
Run status group 3 (all jobs):
WRITE: bw=152MiB/s (159MB/s), 152MiB/s-152MiB/s (159MB/s-159MB/s), io=8192MiB (8590MB), run=53974-53974msec
RHEL 9:
- Encryption type: aes256-cts-hmac-sha384-192
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 0 (all jobs):
WRITE: bw=198MiB/s (208MB/s), 198MiB/s-198MiB/s (208MB/s-208MB/s), io=8192MiB (8590MB), run=41394-41394msec
Run status group 1 (all jobs):
WRITE: bw=204MiB/s (214MB/s), 204MiB/s-204MiB/s (214MB/s-214MB/s), io=8192MiB (8590MB), run=40156-40156msec
Run status group 2 (all jobs):
WRITE: bw=179MiB/s (187MB/s), 179MiB/s-179MiB/s (187MB/s-187MB/s), io=8192MiB (8590MB), run=45868-45868msec
Run status group 3 (all jobs):
WRITE: bw=151MiB/s (158MB/s), 151MiB/s-151MiB/s (158MB/s-158MB/s), io=8192MiB (8590MB), run=54363-54363msec
Overview
| OS | Filesystem | SEC | RANDOM WRITE 8k |
|---|---|---|---|
| Fedora 40 | XFS | – | 2586MiB/s |
| RHEL 7 | NFS | sys | 179MiB/s |
| RHEL 7 | NFS | krb5 | 173MiB/s |
| RHEL 7 | NFS | krb5i | 130MiB/s |
| RHEL 7 | NFS | krb5p | 81.9MiB/s |
| RHEL 8 | NFS | sys | 205MiB/s |
| RHEL 8 | NFS | krb5 | 176MiB/s |
| RHEL 8 | NFS | krb5i | 154MiB/s |
| RHEL 8 | NFS | krb5p | 152MiB/s |
| RHEL 9 | NFS | sys | 189MiB/s |
| RHEL 9 | NFS | krb5 | 182MiB/s |
| RHEL 9 | NFS | krb5i | 180MiB/s |
| RHEL 9 | NFS | krb5p | 162MiB/s |
Storage
Targetcli
Scripts for creating loopback disks
targetcli /loopback create wwn=naa.5000000000000000 \
for i in {00..16}; \
do lvcreate -y -n disk$i -L100G export; \
targetcli /backstores/block create dev=/dev/export/disk$i name=disk$i; \
targetcli /backstores/block/disk$i set attribute optimal_sectors=4096; \
targetcli /loopback/naa.5000000000000000/luns create \
storage_object=/backstores/block/disk$i; \
done
Input Oputput I/
Breef overview of current (2023) I/O mechanism supported Linux kernel.
Synchronous
Represented by system calls read()/write(), pread()/pwrite() and it’s vectored version of readv()/writev() and preadv()/pwritev().
Asynchronous
Represented either by POSIX aio (man aio), libaio and latest io_uring.
Buffered
All IO which ends up in page cache and are not directly written to but are written by the page background process.
Direct I/O:
With direct I/O data is read from or written to the storage device (e.g., HDD, SSD, NVMe) without being cached in the operating system’s buffer cache. This means that data is transferred directly between the application’s memory and the storage device, bypassing any intermediate caching layers in the operating system.
Benefits of Direct I/O:
-
Direct I/O is often used in scenarios where data consistency and control over I/O operations are critical, such as in databases, file systems, and some scientific computing applications.
-
It can help ensure that data is written to and read from the storage device without interference from the operating system’s cache, which can be especially important for applications that require strict data durability or real-time performance. By bypassing the cache, direct I/O can reduce the variability in I/O response times that can occur with cached I/O.
x86_64 REGISTERS
General-Purpose Registers
The 64-bit versions of the ‘original’ x86 registers are named:
- rax - register a extended
- rbx - register b extended
- rcx - register c extended
- rdx - register d extended
- rbp - register base pointer (start of stack)
- rsp - register stack pointer (current location in stack, growing downwards)
- rsi - register source index (source for data copies)
- rdi - register destination index (destination for data copies)
The registers added for 64-bit mode are named:
- r8 - register 8
- r9 - register 9
- r10 - register 10
- r11 - register 11
- r12 - register 12
- r13 - register 13
- r14 - register 14
- r15 - register 15
These may be accessed as:
- 64-bit registers using the r prefix:
rax,r15. - 32-bit registers using the e prefix or d suffix:
eax,r15d. - 16-bit registers using no prefix or a w suffix :
ax,r15w. - 8-bit registers using h (“high byte” of 16 bits) suffix:
ah,bh. - 8-bit registers using l (“low byte” of 16 bits) suffix or ‘b’ suffix:
al,r15b.
================ rax (64 bits)
======== eax (32 bits)
==== ax (16 bits)
== ah (8 bits)
== al (8 bits)
Usage during syscall/function call:
First six arguments are in rdi, rsi, rdx, rcx, r8d, r9d; remaining arguments are on the stack.
For syscalls, the syscall number is in rax. For procedure calls, rax should be set to 0. The called
routine is expected to preserve rsp, rbp, rbx, r12, r13, r14 and r15 but may trampleany other
registers. Return value is in rax.
GIT
Github notes.
Git hub submit changes to pull request:
- rebase
- force push
More verbose:
- changes
- commit,
- git rebase -i HEAD~2’ use ‘f’ to squash the new code into the previous one + the previous commit message
The 30-Year-Old Bug: How Modern Compiler Optimizations Broke a 1994 PRNG
A Deep-Dive Compiler Forensics Case Study
1. The Context: RPM Default Flags and xfstests
The investigation began with compiling xfstests-dev, a standard Linux filesystem testing suite, using Red Hat’s default RPM build optimizations (%optflags).
Modern RPM builds inject a massive payload of performance and security flags, including:
-O2 -flto=auto -ffat-lto-objects -fexceptions -g -grecord-gcc-switches -pipe -Wall -Werror=format-security -Wp,-D_FORTIFY_SOURCE=3 -fstack-protector-strong -m64 -march=x86-64-v3 -mtune=generic ...
When building an Autotools project, these flags must be passed to the ./configure script so they are properly tested and baked into the resulting Makefile, rather than passing them directly to make (which overwrites internal project include flags like -I).
# The correct way to inject RPM flags into Autotools
./configure CFLAGS="$(rpm --eval '%{optflags}')" CXXFLAGS="$(rpm --eval '%{optflags}')" LDFLAGS="$(rpm --eval '%{__global_ldflags}')"
make
2. The Mystery: Diverging Execution
Once compiled, a bizarre issue emerged in the nametest binary. Running the exact same code with the exact same pseudo-random number generator (PRNG) seed (-s 1) produced slightly different outputs depending on the compilation flags:
Unstripped (Standard Flags):
creates: 10 OK, 0 EEXIST (10 total, 0% EEXIST)
Stripped (RPM Flags with LTO & -O2):
creates: 10 OK, 1 EEXIST (11 total, 9% EEXIST)
Because the threshold for these operations relied on op = random() % 100, this divergence proved that the internal state of the PRNG was generating different mathematical sequences despite starting with the identical seed.
3. The Red Herrings
Debugging compiler differences is notoriously difficult. Several theories were tested and ultimately ruled out:
-
Implicit Function Declarations (GCC 14): Did the strictness of GCC 14 cause
./configureto fail a feature check, silently falling back to glibc’srand()instead of POSIXrandom()?- Ruled Out:
nmoutput proved xfstests was statically compiling its own customrandom.cimplementation.
- Ruled Out:
-
charSignedness: Red Hat RPM flags include-funsigned-char. Did casting bytes from a signed array alter the PRNG state?- Ruled Out: Recompiling with
-fsigned-chardid not fix the issue.
- Ruled Out: Recompiling with
-
Missing
-fwrapv: Was signed integer overflow causing the optimizer to alter the math?- Initial Test Failed: Passing
-fwrapvtoCFLAGSdid not fix the issue. (Note: This was a false negative due to missingLDFLAGS, which became critical later.)
- Initial Test Failed: Passing
4. The Smoking Gun: Assembly Analysis
To bypass the compiler’s “black box,” the raw assembly of both binaries was dumped and compared:
objdump -d --no-show-raw-insn -M intel ./nametest-stripped > stripped.asm
objdump -d --no-show-raw-insn -M intel ./nametest-unstripped > unstripped.asm
The C code in nametest.c contained two back-to-back PRNG calls inside a loop:
ip = &table[ random() % totalnames ];
op = random() % 100;
Because of Link-Time Optimization (-flto), GCC merged random.c and nametest.c, allowing it to inline the PRNG math directly into the loop.
Looking at stripped.asm, the optimizer did something highly destructive:
1940: mov edi,DWORD PTR [r12] # edi = original_it
# ...
1777: lea eax,[rdi*4+0x0] # new_it = original_it * 4 <--- THE SMOKING GUN
177e: lea edx,[rax-0x1]
1781: mov DWORD PTR [r12],eax # save new_it to memory
The compiler had hardcoded original_it * 4 to handle both PRNG calls simultaneously, completely skipping a critical safety branch inside the random.c logic.
5. The Root Cause: Weaponized Undefined Behavior
The legacy 1994 PRNG code in random.c contained the following logic:
if (it <= 0)
it = (it + it) ^ MASK;
else
it = it + it;
The original author explicitly relied on signed integer overflow. When it + it exceeded the 32-bit integer limit, it would wrap around and become a negative number. The next time the function was called, the if (it <= 0) branch would catch the negative number and apply the MASK.
How Modern GCC Handled It:
In the C standard, signed integer overflow is Undefined Behavior (UB).
When LTO inlined both random() calls, GCC’s value-range analysis looked at the second call and assumed: “If it was positive in the first call, it + it cannot mathematically be negative because signed overflow is illegal. Therefore, the < 0 branch is impossible code.”
GCC completely deleted the safety branch and simply multiplied the variable by 4. Once the seed reached 1,073,741,824, it * 4 overflowed into 0. The MASK was missed, permanently altering the mathematical sequence of the PRNG for the rest of the test.
6. The Fix
The correct way to preserve 1994 logic in a 2024 compiler pipeline is to remove the Undefined Behavior entirely. In C, unsigned integer overflow is 100% legal and defined to wrap modulo $2^n$.
By casting the variables to unsigned integers for the arithmetic, the compiler is forced to respect the wrap-around, preserving the original signed check without triggering the optimizer’s UB deletion.
The Patched Code (random.c):
it = is[0];
leh = is[1];
/* Use unsigned arithmetic to safely wrap the overflow without triggering UB */
uint32_t u_it = (uint32_t)it;
if (it <= 0)
u_it = (u_it + u_it) ^ MASK;
else
u_it = u_it + u_it;
it = (int32_t)u_it;
With this patch, the PRNG safely survives Link-Time Optimization (-flto) and aggressive -O2 heuristics, ensuring the stripped and unstripped binaries generate the exact same random sequence.
7. Standalone Reproduction
The bug was reproduced in isolation to confirm the root cause without involving xfstests. The reproduction lives in:
randtest/
├── random.c — PRNG implementation (_irandm, _random, random, srandom, get_it)
├── random.h — declarations
├── randtest.c — main: two sequential random() calls per iteration with diagnostics
└── bug_demo.c — self-contained single-file reproduction (see below)
7.1 Two-file reproduction (random.c + randtest.c)
randtest.c calls random() twice per iteration and reads saved_seed[0] via get_it() before and after each call, so the corrupted PRNG state is visible directly:
# correct behaviour
gcc -O0 -o randtest_plain randtest.c random.c
# triggers the bug via LTO cross-file inlining
gcc $(rpm --eval '%{optflags}') $(rpm --eval '%{build_ldflags}') \
-o randtest_rpm randtest.c random.c
Output diverges at iteration 16 — the first iteration after it crosses 2^30 and the doubled value wraps past INT32_MAX:
# plain (correct)
iter 16: it=1187941550 r1=138120522 it=-1919084196 r2=1734299348 it=945648879 <-- OVERFLOW
# RPM (buggy)
iter 16: it=1187941550 r1=138120522 it=-1919084196 r2=2084125751 it=456798904
r1 is still identical (both calls share the same entry state for that iteration), but r2 diverges because the inlined second call skips the MASK branch and writes a corrupted value to saved_seed[0].
7.2 Single-file reproduction (bug_demo.c)
The bug does not require LTO. A single translation unit compiled with plain -O2 is enough — the compiler inlines _irandm freely within the file and applies the same cross-call value-range analysis.
gcc -O2 -fno-lto -o bug_demo_O2 bug_demo.c # triggers the bug
gcc -O0 -o bug_demo_O0 bug_demo.c # correct reference
Same divergence, same iteration, no linker flags involved.
7.3 Why the printf inside _irandm masks the bug
During development a diagnostic printf was placed inside _irandm itself. This silently cured the bug: printf is an external call with side effects, which GCC treats as a full memory barrier. The optimizer can no longer track it across the call boundary, so it conservatively keeps both branches. The numbers stayed identical across builds.
The only visible artifact was that the RPM build silently dropped the "<-- OVERFLOW" label from its output — the ternary (it_old > 0 && it < 0) was statically eliminated because GCC knew (from UB reasoning in the else branch) that it < 0 is impossible after it = it + it with a positive it. Correct numbers, missing label: a subtler manifestation of the same UB exploitation.
Moving the printf to main (via get_it()) removed the barrier and restored the divergence.
8. Assembly Deep-Dive (bug_demo_O2.s)
The generated assembly (bug_demo_O2.s) shows the inlined loop. The critical section (annotated):
.L7: ; loop top — load saved_seed
leal (%rdx,%rdx), %r8d ; r8d = it + it (call 1 new_it; may wrap!)
testl %edx, %edx ; test ORIGINAL it (before doubling)
jg .L2 ; it > 0 → fast path, NO branch check for call 2
xorl $593970775, %r8d ; MASK applied (call 1, it ≤ 0 path)
...
testl %r8d, %r8d ; call 2 gets its own branch — but only from
jg .L4 ; the it ≤ 0 entry path
xorl $593970775, %esi ; MASK for call 2 if needed
.L2: ; entered when original it was positive
leal -1(%r8), %esi ; nit1 = (it+it) - 1
andl $127, %esi
imull mt(,%rsi,4), %eax ; leh *= mt[nit1 & 127]
leal 0(,%rdx,4), %esi ; THE BUG: esi = it * 4
; GCC assumed it+it > 0 (signed overflow = UB)
; so it*2 again needs no sign check.
; When it = 2^30: it*4 = 2^32 → truncates to 0.
...
; falls straight into call 2 with esi = it*4, no branch, no MASK
The testl %edx, %edx / jg .L2 pair is the only branch serving both calls. It tests the pre-doubling it, not the result of the doubling (%r8d). Once jg .L2 is taken, call 2’s if (it <= 0) check is gone entirely — replaced by the hardcoded leal 0(,%rdx,4).
In contrast, -O0 emits two independent call _random instructions. Each is a black box; no cross-call value-range analysis is possible and _irandm executes its branch correctly on every invocation.
Trigger conditions summary
| Scenario | Bug triggers |
|---|---|
Separate files, -O0 | No — no inlining |
Separate files, -O2, no LTO | No — compiler cannot see across files |
Separate files, -O2, -flto | Yes — LTO merges IR, inlines across files |
Single file (bug_demo.c), -O2, no LTO | Yes — inlined within one translation unit |
Single file, -O0 | No — no inlining, no value-range optimization |
Any flags, printf inside _irandm | No — printf acts as a memory barrier |
9. Compiler Options Reference
Flags that trigger the bug
| Flag | Role |
|---|---|
-O2 | Enables inlining and value-range propagation — the minimum level needed |
-flto=auto | Link-Time Optimization: merges all translation units into one IR before optimization, giving the same cross-file view as a single .c file |
-ffat-lto-objects | Embeds both LTO IR and regular object code in each .o, so the archive works with and without LTO-aware linkers |
Flags in %{optflags} relevant to this bug
-O2 -flto=auto -ffat-lto-objects -fexceptions -g -grecord-gcc-switches -pipe
-Wall -Wno-complain-wrong-lang -Werror=format-security
-Wp,-U_FORTIFY_SOURCE,-D_FORTIFY_SOURCE=3 -Wp,-D_GLIBCXX_ASSERTIONS
-specs=/usr/lib/rpm/redhat/redhat-hardened-cc1
-fstack-protector-strong
-specs=/usr/lib/rpm/redhat/redhat-annobin-cc1
-m64 -march=x86-64 -mtune=generic
-fasynchronous-unwind-tables -fstack-clash-protection
-fcf-protection -mtls-dialect=gnu2
-fno-omit-frame-pointer -mno-omit-leaf-frame-pointer
Of these, -O2 and -flto=auto are the two flags directly responsible for the bug. The rest add security hardening and debug info but do not influence the PRNG optimization.
Flags in %{build_ldflags} relevant to this bug
-Wl,-z,relro -Wl,--as-needed -Wl,-z,pack-relative-relocs -Wl,-z,now
-specs=/usr/lib/rpm/redhat/redhat-hardened-ld
-specs=/usr/lib/rpm/redhat/redhat-hardened-ld-errors
-specs=/usr/lib/rpm/redhat/redhat-annobin-cc1
-Wl,--build-id=sha1
The hardened-ld specs activate the LTO linker plugin. Without passing LDFLAGS correctly, -flto in CFLAGS compiles LTO IR into the objects but the link step discards it — which is why early -fwrapv tests appeared to fix the issue (the LTO IR was never linked).
Flags that suppress the bug
| Flag | Effect |
|---|---|
-O0 | Disables inlining entirely; each _irandm call is a real function call |
-fno-lto | Disables LTO; cross-file inlining impossible |
-fwrapv | Tells GCC that signed integer overflow wraps (two’s complement); UB assumption removed, branch preserved |
-fno-strict-overflow | Weaker form of -fwrapv; disables overflow-based optimizations |
__attribute__((noinline)) on _irandm | Prevents inlining of the function; forces a real call boundary |
Security hardening visible in the RPM binary (unrelated to the bug)
| Feature | Flag | Effect on binary |
|---|---|---|
| PIE | -specs=redhat-hardened-cc1 | ELF type changes from EXEC to DYN |
| Full RELRO | -Wl,-z,relro + BIND_NOW | All GOT entries made read-only after startup |
| BIND_NOW | -Wl,-z,now | All symbols resolved at load time |
| FORTIFY_SOURCE=3 | -D_FORTIFY_SOURCE=3 | printf replaced by __printf_chk, buffer overflows detected at runtime |
| Stack clash protection | -fstack-clash-protection | Probe stack pages on allocation to prevent stack-clash attacks |
| CF protection (CET) | -fcf-protection | endbr64 inserted at every indirect-jump target |
| Frame pointers kept | -fno-omit-frame-pointer | Enables reliable stack unwinding in profilers and crash dumps |
| Debug info | -g -grecord-gcc-switches | DWARF sections embedded; binary grows from 13K to 19K |
2’s Complement
The dominant way modern hardware represents signed integers.
Core Idea
For an N-bit integer, a negative number -x is stored as 2^N - x.
For 8-bit (N=8):
-1 → 2^8 - 1 = 255 = 0xFF = 11111111
-2 → 2^8 - 2 = 254 = 0xFE = 11111110
-128 → 2^8 - 128 = 128 = 0x80 = 10000000
Bit Layout (8-bit)
Bit pattern | Unsigned | Signed (2's complement)
-------------|----------|------------------------
0000 0000 | 0 | 0
0000 0001 | 1 | 1
0111 1111 | 127 | 127 ← INT_MAX
1000 0000 | 128 | -128 ← INT_MIN (sign bit flips)
1000 0001 | 129 | -127
1111 1110 | 254 | -2
1111 1111 | 255 | -1
The sign bit (MSB) being 1 means negative. The range is asymmetric: one more negative value than positive.
How to Negate Manually
Two equivalent methods:
Method 1: Flip all bits, then add 1
5 = 0000 0101
~5 = 1111 1010 (flip)
-5 = 1111 1011 (add 1)
Method 2: 2^N - x
-5 = 256 - 5 = 251 = 1111 0101 ✓ (same result)
Why Hardware Loves It
Addition and subtraction use the same circuit for signed and unsigned:
0000 0101 (+5)
+ 1111 1011 (-5 in 2's complement)
-----------
1 0000 0000 → carry discarded → 0000 0000 = 0 ✓
No special subtraction hardware needed. This is the primary reason 2’s complement won over alternatives like sign-magnitude or 1’s complement.
The Three Historical Alternatives (Mostly Dead)
| Scheme | How -5 looks (8-bit) | Problem |
|---|---|---|
| Sign-magnitude | 1000 0101 | Two zeros (+0 and -0), complex arithmetic |
| 1’s complement | 1111 1010 | Also two zeros, end-around carry needed |
| 2’s complement | 1111 1011 | One zero, simple arithmetic ✓ |
Reading a Bit Pattern
Example: 1111 1111
Method 1 — sign bit formula (MSB has weight -2^(N-1), rest are normal):
1111 1111
│└──────┘
│ positional values: 64+32+16+8+4+2+1 = 127
│
└─ sign bit: -128
Total: -128 + 127 = -1
Method 2 — flip and add 1:
1111 1111 → flip → 0000 0000 → add 1 → 0000 0001 = 1
Magnitude is 1, sign bit is 1, so the value is -1.
Verify:
1111 1111 (-1)
+ 0000 0001 (+1)
-----------
1 0000 0000 → carry dropped → 0 ✓
All-ones is always
-1in 2’s complement, regardless of bit width (8, 16, 32, 64).
Example: 1000 0000
Method 1 — sign bit formula:
1000 0000
│└──────┘
│ positional values: 0+0+0+0+0+0+0 = 0
│
└─ sign bit: -128
Total: -128 + 0 = -128
Method 2 — flip and add 1:
1000 0000 → flip → 0111 1111 → add 1 → 1000 0000
You get 1000 0000 back — it’s its own negation. This is why -128 has no positive counterpart in 8-bit signed: +128 doesn’t fit (0111 1111 = 127 is the max).
The asymmetry:
INT_MIN = -128 = 1000 0000
INT_MAX = +127 = 0111 1111
|INT_MIN| > INT_MAX — this is why abs(INT_MIN) is undefined behavior in C. Negating -128 would require +128, which overflows.
Same Bits, Different Meaning
1000 0000 represents different values depending on interpretation:
| Interpretation | Value |
|---|---|
| Unsigned | 128 |
| 2’s complement signed | -128 |
The hardware stores 1000 0000 — whether that’s 128 or -128 is decided purely by how your code declares the variable:
uint8_t u = 0x80; // 128
int8_t s = 0x80; // -128
Casting uint32_t to int32_t
When you cast uint32_t v to int32_t:
- The bit pattern does not change
- The CPU just reinterprets bit 31 as a sign bit
- If bit 31 is
1(value ≥2^31), the result is negative
uint32_t v = 0x80000000; // 2147483648, bit 31 set
int32_t s = (int32_t)v; // -2147483648 — same bits, signed interpretation
In C11/C17 this is implementation-defined behavior. In C23 it is finally guaranteed by the standard, as 2’s complement is now mandated for all signed integer types.
Overflow Wraps the Number Line into a Circle
Visualize it as a clock:
0
-1 1
-2 2
...
-128 127 (8-bit)
-127
Adding past INT_MAX wraps to INT_MIN, and vice versa:
result = (a + b) mod 2^N (then reinterpret as signed)
Ansible
VM Provisioning Guide
Docs
Nextcloud internals
Cleaning of the bruteforce IP
use nextcloud;
show tables;
select * from oc_bruteforce_attempts;
delete from oc_bruteforce_attempts where IP="xxx.xxx.xxx.xxx";