Introduction

News

Sailing Notes

Notes

Food

Places

Croatia

North

Central

South

⛵ Sailing Food Plan & Shopping List

Duration: 6 Days | Crew: 8 People

📅 Daily Meal Schedule

DayBreakfastLunch / SaladMain Dinner
Day 1 (Sun)Daily offer / what’s foundPasta with Roasted Pepper PestoRed Lentil Soup
Day 2 (Mon)Daily offer / what’s foundGreek Baked ChickpeasHarira Soup
Day 3 (Tue)Daily offer / what’s foundBlack Lentil SaladVegetable Risotto
Day 4 (Wed)Daily offer / what’s foundPoke BowlChili sin Carne
Day 5 (Thu)Daily offer / what’s foundShakshuka / LečoButter Beans
Day 6 (Fri)Daily offer / what’s foundVeggie CouscousLeftover Feast

See Shopping List for the full ingredient list and water requirements.


⚓ Yacht Cooking Tips

  1. Freshness Order: Use mushrooms, spinach, and soft tomatoes first. Cabbage, carrots, and potatoes can wait until the end of the week.
  2. Bread Management: Start with fresh bakery bread. Switch to vacuum-packed “bake-off” bread that you can toast in the oven from Day 4 onwards.
  3. Storage: Heavy items (water, cans, potatoes) should be stored low in the boat (the bilge) to keep the boat stable.
  4. Cleaning: Save fresh water by rinsing dishes in seawater first, then doing a final quick rinse with fresh water.

🍲 Recipes

🛒 Master Shopping List

Calculated for 8 People | 6 Days


🥬 Fresh Produce

ItemQuantityUsed in
Onions (yellow)~15 pcs (~3 kg)Red Lentil Soup, Greek Baked Chickpeas, Harira, Risotto, Chili, Shakshuka, Couscous
Onion (red)6Black Lentil Salad
Garlic5 headsAll savoury dishes
Leeks7 large stalksRed Lentil Soup, Greek Baked Chickpeas
Carrots~2 kg (~15 pcs)Red Lentil Soup, Harira, Risotto, Poke Bowl
Bell peppers20 pcs (mix of colors)Chili sin Carne, Shakshuka
Zucchini6Vegetable Risotto, Veggie Couscous
Fresh tomatoes14Veggie Couscous (4), Butter Beans (4)
Cucumber8Poke Bowl
Avocado8 (mix ripe/firm)Poke Bowl
Celery2 large bunches (~12 stalks)Red Lentil Soup, Harira
Fresh parsley3 bunchesBlack Lentil Salad, Shakshuka, Pasta with Tuna, Veggie Couscous
Fresh cilantro2 bunchesHarira, Veggie Couscous
Lemons6Red Lentil Soup, Harira

🧀 Fridge & Proteins

ItemQuantityUsed in
Eggs30Shakshuka + breakfasts
Klobása / hard salami400 gGreek Baked Chickpeas
Tofu (firm)3 packs (~1.2 kg)Greek Baked Chickpeas (400 g) + vegan Shakshuka option (800 g)
Parmesan400 gPasta with Pepper Pesto (80 g), Red Lentil Soup (rinds + garnish), Risotto (100 g)
Mozzarella4 ballsBlack Lentil Salad
Smoked salmon800 gPoke Bowl
Wakame (frozen)500 gPoke Bowl
Hard cheese (Eidam / Gouda)1 kgBreakfasts & snacking
Fresh spinach400 gButter Beans
Vegan heavy cream (oat/soy)4 packs (200 ml each)Butter Beans

🥫 Pantry & Dry Goods

ItemQuantityUsed in
Pasta (penne / spaghetti)2 kgPasta with Pepper Pesto (1.5 kg) + extra
Arborio rice1 kgVegetable Risotto
Sushi rice1 kgPoke Bowl
Couscous1 kgVeggie Couscous
Buckwheat3 kgBreakfasts
Vermicelli / soup noodles100 gHarira
Canned crushed tomatoes10 cans (400 g)Red Lentil Soup (2), Harira (2), Chili (2), Shakshuka (4)
Roasted red peppers (jars)4 jars (460 g each)Pasta with Pepper Pesto
Tomato paste2 tubes (200 g each)Harira, Chili, Veggie Couscous
Canned chickpeas8 cans (310 g each)Greek Baked Chickpeas (4), Harira (4)
Canned beans (kidney or black)4 cans (400 g each)Chili sin Carne
Canned corn6 cans (340 g each)Vegetable Risotto (2), Chili sin Carne (2)
Canned peas6 cans (400 g each)Vegetable Risotto (2), Veggie Couscous (2)
Canned mushrooms6 cans (400 g each)Vegetable Risotto (2), Veggie Couscous (2)
Canned white beans8 cans (400 g each)Butter Beans
Lentils (black / beluga)500 gBlack Lentil Salad
Lentils (red)1 kgRed Lentil Soup (~500 g), Harira (200 g)
Olives (pitted)1 large jar (900 g)Black Lentil Salad
Capers1 jar (100 g)Black Lentil Salad
Sun-dried tomatoes (in oil)2 jars (280 g each)Butter Beans
Nori / seaweed2 packs (25 g each)Poke Bowl
Kimchi2 large jars (500 g each)Poke Bowl
Edamame (canned)2 cans (400 g each)Poke Bowl
Bamboo slices (canned)2 cans (540 g each, AROY-D)Poke Bowl

🧂 Spices & Essentials

ItemQuantityUsed in
Olive oil2 LAll savoury dishes
Soy sauce1 bottlePoke Bowl
Teriyaki sauce1 bottlePoke Bowl
Rice vinegar1 bottlePoke Bowl
Red wine vinegar1 bottleBlack Lentil Salad
Sriracha1 bottlePoke Bowl
Wasabi1 small packPoke Bowl
Sesame seeds1 small packPoke Bowl
Harissa paste1 jarHarira
Cumin1 jarHarira, Chili, Shakshuka, Veggie Couscous
Smoked paprika1 jarHarira, Chili, Shakshuka, Veggie Couscous
Cinnamon1 jarHarira
Turmeric1 jarHarira
Chili flakes1 jarChili sin Carne
Bay leaves1 packRed Lentil Soup
Dried sage or rosemary1 jarGreek Baked Chickpeas
Dark chocolate1 barChili sin Carne (1 square)
Sugar1 packPoke Bowl rice seasoning
Salt1 kgAll dishes
Black pepper1 grinderAll dishes
Vegetable bouillon3 packs (cubes/powder)~8 L broth across Red Lentil Soup, Harira, Risotto, Couscous
Bread12 loaves fresh (Days 1–3) + 6 packs bake-off baguettes (Days 4–6)Serving with Greek Baked Chickpeas, Harira, Shakshuka
Snacks10 packs biscuits / crackers, 5 bags nuts, chips—
Trash bags1 large pack—
Baking paper1 rollGreek Baked Chickpeas
Aluminium foil1 rollGreek Baked Chickpeas

🍺 Drinks

  • Beer: 6 cans/person/day min
  • Port wine
  • Rum: 2 bottles min
  • Coke + juice

💧 Water Requirements (3 L per person/day)

To ensure a safe supply for 8 people for 8 days (approx. 200 Liters):

  • Option 1 (1.5 L Bottles): ~135 bottles (approx. 23 packs of six)
  • Option 2 (2 L Bottles): ~100 bottles (approx. 17 packs of six)

Tip: Bring a permanent marker to label bottle caps with names to avoid waste.

Pasta with Roasted Pepper Pesto

Serves: 8 | Prep: 10 min | Cook: 15 min

Ingredients

  • 1.5 kg pasta (penne or fusilli)
  • 4 large jars roasted red peppers (or 8 fresh peppers, roasted and peeled)
  • 100 ml olive oil
  • 4 garlic cloves
  • 80 g parmesan, grated
  • Salt & pepper

Instructions

  1. Blend roasted peppers, olive oil, garlic, and parmesan until smooth. Season.
  2. Cook pasta al dente. Reserve 1 cup pasta water before draining.
  3. Toss pasta with pesto, loosening with pasta water as needed. Serve with extra parmesan.

Red Lentil Soup

Serves: 5 | Prep: 10 min | Cook: 50 min

Ingredients

  • 4 tbsp extra virgin olive oil
  • 2 medium yellow onions, medium dice
  • 2 medium leeks, white and pale green parts only, rinsed and medium dice
  • 3 medium carrots (~½ lb), medium dice
  • 5 celery stalks (~6 oz), medium dice
  • 2 large garlic cloves, roughly chopped
  • 1½ cups red split lentils, rinsed
  • 1 can (15½ oz) crushed Italian tomatoes
  • 1–2 parmigiano-reggiano rinds
  • 2 dried bay leaves
  • 2 quarts (8 cups) low-sodium chicken or vegetable broth
  • Kosher salt & freshly ground black pepper
  • Freshly grated parmigiano-reggiano, for serving (optional)

Instructions

  1. Heat olive oil in a large soup pot over medium heat. Add onions, leeks, and a generous pinch of salt. Sauté with lid askew, stirring occasionally, until soft and translucent — about 10–15 minutes.
  2. Add carrots, celery, and garlic; stir for 3–4 minutes. Add lentils, crushed tomatoes, parmesan rinds, bay leaves, and broth. Bring to a boil, then reduce to medium-low and simmer uncovered for 30–40 minutes, stirring every 10 minutes, until vegetables are tender and lentils have broken down. Soup should be thick and hearty.
  3. Optionally blend a small portion with an immersion blender for a better texture.
  4. Season with salt and pepper. Add a squeeze of lemon if it tastes flat. Remove bay leaves and parmesan rinds before serving. Garnish with grated parmesan.

Greek Baked Chickpeas

Serves: 8 | Prep: 15 min | Cook: 60 min

Ingredients

  • 4 cans chickpeas, drained
  • 4 large leeks, white and pale green parts, sliced into rounds
  • 2 onions, sliced
  • 400 g tofu, cubed
  • 400 g klobása / hard salami, sliced
  • 6 tbsp olive oil
  • 4 garlic cloves, minced
  • 1 tbsp dried sage or rosemary
  • 300 ml water or vegetable broth
  • Salt & pepper
  • Bread for serving

Instructions

  1. Preheat oven to 180°C.
  2. Combine chickpeas, leeks, onion, tofu, klobása, garlic, olive oil, and sage in a large baking dish. Add water and season generously.
  3. Cover with foil and bake for 45 minutes. Remove foil and bake a further 15–20 minutes until golden.
  4. Serve with crusty bread.

Harira Soup

Serves: 8 | Prep: 15 min | Cook: 40 min

Moroccan tomato soup with chickpeas and lentils — warming, spiced, and thick.

Ingredients

  • 4 cans chickpeas, drained
  • 200 g red lentils
  • 2 onions, diced
  • 4 celery stalks, diced
  • 2 carrots, diced
  • 2 cans crushed tomatoes
  • 2 tbsp tomato paste
  • 2 tbsp harissa paste
  • 1 tsp turmeric
  • 1 tsp cumin
  • ½ tsp cinnamon
  • 1 tsp smoked paprika
  • 100 g vermicelli
  • 2 L vegetable broth
  • Olive oil, salt & pepper
  • Fresh parsley & cilantro
  • Lemon wedges for serving

Instructions

  1. Sauté onion, celery, and carrots in olive oil for 5 minutes until softened.
  2. Add garlic and spices; cook 1 minute.
  3. Add crushed tomatoes, tomato paste, lentils, chickpeas, and broth. Bring to a boil then simmer uncovered for 25 minutes.
  4. Add vermicelli and cook 8 minutes more. Stir in harissa.
  5. Season and serve with lemon wedges, fresh parsley, and bread.

Black Lentil Salad

Serves: 8 | Prep: 10 min | Cook: 25 min

Ingredients

  • 500 g black / beluga lentils
  • 1 red onion, finely diced
  • 3 tbsp capers
  • 200 g olives, pitted and halved
  • 4 mozzarella balls, torn
  • 4 tbsp olive oil
  • 2 tbsp red wine vinegar
  • Salt & pepper
  • Fresh parsley

Instructions

  1. Cook lentils in salted boiling water for 20–25 minutes until tender. Drain and cool slightly.
  2. While still warm, toss with olive oil, vinegar, salt, and pepper.
  3. Mix in red onion, capers, and olives.
  4. Top with torn mozzarella and fresh parsley. Serve at room temperature.

Vegetable Risotto

Serves: 8 | Prep: 15 min | Cook: 35 min

Ingredients

  • 1 kg arborio rice
  • 1 kg mushrooms, sliced
  • 2 zucchini, diced
  • 2 cans peas, drained
  • 2 cans corn, drained
  • 2 carrots, diced
  • 2 onions, diced
  • 4 garlic cloves, minced
  • 2 L vegetable broth, kept hot
  • 100 g parmesan, grated
  • Olive oil, salt & pepper

Instructions

  1. Sauté onion and garlic in olive oil until soft.
  2. Add mushrooms and carrots; cook until softened, about 8 minutes.
  3. Add rice and stir to coat. Pour in a ladle of hot broth; stir until absorbed. Repeat, adding broth one ladle at a time, for 18–20 minutes.
  4. Stir in zucchini, peas, and corn in the last 5 minutes.
  5. Finish with a drizzle of olive oil and parmesan. Season and serve immediately.

Poke Bowl

Serves: 8 | Prep: 20 min | Cook: 20 min

Ingredients

  • 1 kg sushi rice
  • 8 avocados, sliced
  • 2 jars kimchi
  • 2 packs nori / seaweed, cut into strips
  • 4 cucumbers, sliced
  • 4 carrots, julienned
  • 4 tbsp soy sauce
  • 2 tbsp sesame seeds
  • Sriracha, wasabi & extra soy sauce for serving

Instructions

  1. Cook sushi rice per package instructions. Season with a splash of rice vinegar, a pinch of sugar, and salt.
  2. Divide rice into bowls.
  3. Arrange avocado, kimchi, seaweed, cucumber, and carrot on top.
  4. Drizzle with soy sauce and sprinkle sesame seeds. Serve with sriracha and wasabi on the side.

Chili sin Carne

Serves: 8 | Prep: 15 min | Cook: 35 min

Ingredients

  • 4 cans kidney or black beans, drained
  • 2 cans corn, drained
  • 2 cans crushed tomatoes
  • 4 bell peppers, diced
  • 2 onions, diced
  • 4 garlic cloves, minced
  • 2 tbsp tomato paste
  • 1 tbsp cumin
  • 1 tbsp smoked paprika
  • 1 tsp chili flakes
  • 1 square dark chocolate
  • Olive oil, salt & pepper

Instructions

  1. Sauté onion, peppers, and garlic in olive oil until soft, about 8 minutes.
  2. Add cumin, paprika, and chili flakes; cook 1 minute.
  3. Add beans, corn, crushed tomatoes, and tomato paste. Stir and simmer 25–30 minutes.
  4. Stir in dark chocolate at the end until melted. Season.
  5. Serve with bread or rice.

Shakshuka / Lečo

Serves: 8 | Prep: 10 min | Cook: 30 min

Eggs (or tofu) poached in spiced tomato and pepper sauce.

Ingredients

  • 8 bell peppers, diced
  • 4 cans crushed tomatoes
  • 2 onions, diced
  • 4 garlic cloves, minced
  • 16 eggs (or 800 g firm tofu, cubed, for vegan)
  • 2 tsp smoked paprika
  • 1 tsp cumin
  • Olive oil, salt & pepper
  • Fresh parsley

Instructions

  1. Sauté onion and garlic in olive oil until soft.
  2. Add peppers; cook 5 minutes.
  3. Add crushed tomatoes and spices; simmer 15 minutes until sauce thickens.
  4. Make wells in the sauce and crack in eggs (or nestle tofu cubes). Cover and cook until eggs are just set, about 8–10 minutes.
  5. Season and garnish with parsley. Serve with bread.

Butter Beans

Serves: 8 | Prep: 10 min | Cook: 25 min

Creamy white bean stew with spinach, tomatoes, and sun-dried tomatoes in a rich vegan cream sauce.

Ingredients

  • 8 cans white beans, drained
  • 4 packs vegan heavy cream (oat or soy, 200 ml each)
  • 400 g fresh spinach (or 500 g frozen)
  • 2 jars sun-dried tomatoes in oil, roughly chopped
  • 4 fresh tomatoes, diced
  • 4 tbsp tomato puree
  • 2 onions, diced
  • 4 garlic cloves, minced
  • Olive oil, salt & pepper

Instructions

  1. Sauté onion in olive oil over medium heat until soft, about 5 minutes. Add garlic and cook 1 minute.
  2. Add diced tomatoes, sun-dried tomatoes, and tomato puree. Simmer 8 minutes until tomatoes break down.
  3. Add white beans and vegan cream. Stir and simmer 10 minutes until sauce thickens.
  4. Fold in spinach and cook until just wilted, 2–3 minutes. Season with salt and pepper.
  5. Serve with crusty bread.

Veggie Couscous

Serves: 8 | Prep: 10 min | Cook: 20 min

Ingredients

  • 1 kg couscous
  • 2 zucchini, diced
  • 2 cans peas, drained
  • 4 tomatoes, diced
  • 2 tbsp tomato paste
  • 2 onions, diced
  • 4 garlic cloves, minced
  • 1 tsp cumin
  • 1 tsp smoked paprika
  • 2 tbsp olive oil
  • ~1 L vegetable broth (hot, for couscous)
  • Fresh parsley or cilantro, salt & pepper

Instructions

  1. Sauté onion and garlic in olive oil until soft.
  2. Add zucchini, cumin, and paprika; cook 5 minutes.
  3. Add diced tomatoes and tomato paste; simmer 10 minutes. Stir in peas. Season.
  4. Place couscous in a large bowl. Pour hot broth over in a 1:1 ratio, cover and rest 5 minutes, then fluff with a fork.
  5. Serve couscous topped with the vegetable sauce. Garnish with fresh parsley.

Porrusalda

Serves: 8 | Prep: 15 min | Cook: 55 min

Traditional Basque leek and potato soup — simple, hearty, and warming.

Ingredients

  • 10 large leeks, white and pale green parts only, sliced into rounds and rinsed well
  • 3 kg potatoes, broken into rough chunks (press with thumb against knife edge — don’t cut cleanly; rough edges release more starch)
  • 3 carrots, diced
  • 6 tbsp olive oil
  • 4 garlic cloves, minced
  • 2 bay leaves
  • 2 L vegetable broth or water
  • Salt & pepper
  • Fresh parsley, crusty bread for serving

Instructions

  1. Sauté leeks in olive oil with a pinch of salt for 8–10 minutes until softened.
  2. Add garlic and carrots; cook 3 minutes.
  3. Add potato chunks, bay leaves, and broth. Bring to a boil.
  4. Reduce heat and simmer covered for 40–50 minutes until potatoes are very tender.
  5. Remove bay leaves. Optionally blend a small portion for a creamier texture.
  6. Season generously. Serve with crusty bread and a drizzle of olive oil.

Pasta with Tuna

Serves: 8 | Prep: 10 min | Cook: 15 min

Ingredients

  • 1.5 kg spaghetti or penne
  • 10 cans tuna, drained
  • 200 g olives, pitted and halved
  • 3 tbsp capers
  • 2 jars sun-dried tomatoes, roughly chopped
  • 4 garlic cloves, minced
  • 4 tbsp olive oil
  • Fresh parsley
  • Chili flakes, salt & pepper

Instructions

  1. Cook pasta al dente. Reserve pasta water before draining.
  2. Warm olive oil and garlic in a large pan over medium heat. Add olives, capers, and sun-dried tomatoes; cook 2 minutes.
  3. Add tuna and toss gently — don’t break it up too much.
  4. Toss with drained pasta, adding a splash of pasta water to bring it together.
  5. Finish with parsley and chili flakes. Season.

Navigation — Staircase Method

The staircase method converts between four course types. The rule is simple:

  • Going up (Compass → Water): ADD
  • Going down (Water → Compass): SUBTRACT

The Staircase

┌──────────────────────────────────────────────────┐
│  KV  — Water Course     (Kurz vůči vodě)         │
│        ↑ + drift         ↓ − snos                │
│  KR  — True Course      (Pravý kurz)             │
│        ↑ + variation     ↓ − variace             │
│  KM  — Magnetic Course  (Magnetický kurz)        │
│        ↑ + deviation     ↓ − deviace             │
│  KK  — Compass Course   (Kompasový kurz)         │
└──────────────────────────────────────────────────┘
SymbolNameCzech
KKCompass CourseKompasový kurz
KMMagnetic CourseMagnetický kurz
KRTrue CoursePravý kurz
KVWater CourseKurz vůči vodě
devDeviationDeviace
varVariationVariace
snosDriftSnos

Signs

  • W (West) var/dev → negative value
  • E (East) var/dev → positive value

Variation (var)

Variation is the angle between True North and Magnetic North. Always update it from the chart for the current year:

Current var = Base var ± (annual change × years elapsed)

Example: Base 5°35’ W in 2018, annual change 6’ W, year 2023:

var = 5°35' + (5 × 6') = 5°35' + 30' = 6°05' W ≈ −6°

KK → KV (Compass to Water — going UP, ADD)

KM = KK + dev
KR = KM + var
KV = KR + snos

Example: KK = 261°, dev = 0°, var = −6°, wind from starboard (snos = −4°)

KM = 261° + 0°   = 261°
KR = 261° + (−6°) = 255°
KV = 255° + (−4°) = 251°

KV → KK (Water to Compass — going DOWN, SUBTRACT)

KR = KV − snos
KM = KR − var
KK = KM − dev

Example: KR = 346°, dev = 0°, var = −6°

KM = 346° − (−6°) = 352°
KK = 352° − 0°    = 352°

Drift (snos)

Drift is the sideways push of the wind on the boat, shifting the water course from the true course.

Wind directiondrift sign
Wind from starboard (right)− (subtract)
Wind from port (left)+ (add)

Bearings

The same staircase applies when converting compass bearings to true bearings for chart plotting:

True bearing = Compass bearing + dev + var

Plot two or more true bearings on the chart — their intersection is your estimated position.


Distance & Speed

Distance (NM) = Speed (kts) × Time (min) / 60
Time (min)    = Distance (NM) × 60 / Speed (kts)

Example: Speed = 3 kts, Time = 90 min

Distance = 3 × 90 / 60 = 4.5 NM

ETA

ETA = departure time + travel time
travel time (min) = Distance × 60 / Speed

Example: Depart 16:00, distance = 9.6 NM, speed = 3 kts

travel time = 9.6 × 60 / 3 = 192 min = 3h 12min
ETA = 16:00 + 3:12 = 19:12

Converting Compass Bearing to Chart Bearing (Position Fix)

When you sight a landmark with a hand bearing compass, the reading is a compass bearing. To plot it on a nautical chart (which uses true north) you must convert it using the same staircase.

Steps

  1. Sight the landmark with the hand bearing compass → get compass bearing
  2. Convert to true bearing (go UP, ADD):
    True Bearing = Compass bearing + dev + var
    
  3. Plot on chart: from the landmark, draw a line in the reciprocal direction (true bearing ± 180°) — you are somewhere along this line
  4. Repeat for a second (and ideally third) landmark
  5. Intersection of the lines = your position fix

Example

Compass bearing to lighthouse = 043°, dev = 0°, var = −6°

True Bearing = 043° + 0° + (−6°) = 037°
Reciprocal   = 037° + 180°        = 217°

Draw a line from the lighthouse at 217° on the chart. You are on that line.

Tips

  • Use 3 bearings when possible — if they form a small triangle (cocked hat), your position is inside it
  • Choose landmarks roughly 60–120° apart for the best intersection angle
  • Bearings to objects ahead or astern are less reliable — prefer objects to the side
  • Take all bearings quickly to minimise error from boat movement

Measuring Compass Deviation

Deviation is caused by the boat’s own magnetic field (engine, wiring, steel fittings). It varies with heading and must be measured for each compass on each boat.

Formula

dev = True Course − var − KK

Where True Course comes from GPS, a transit, or a known bearing.


Method 1: GPS Comparison (easiest)

Modern GPS units show COG (Course Over Ground) in true degrees.

Requirements: calm water, no current, no wind (so the boat tracks straight — leeway = 0).

Steps:

  1. Motor (not sail — sails cause leeway) on a steady heading
  2. Note the compass heading (KK)
  3. Read the GPS COG (= True Course)
  4. Calculate deviation:
    dev = GPS COG − var − KK
    
  5. Repeat on at least 8 headings (N, NE, E, SE, S, SW, W, NW)
  6. Record results in a deviation table

Example: KK = 090°, var = −6°, GPS COG = 087°

dev = 087° − (−6°) − 090° = 087° + 6° − 090° = +3°E

Method 2: Transit / Leading Lines (most accurate)

A transit is two charted objects that line up — their true bearing is known from the chart.

Steps:

  1. Find two charted objects that form a transit (e.g. lighthouse and church spire)
  2. Look up or calculate the true bearing of the transit from the chart
  3. Motor until both objects are exactly in line
  4. Note the compass heading at that moment
  5. Calculate:
    dev = True Bearing − var − Compass reading
    

Advantage: No GPS needed, very precise.


Method 3: Reciprocal Bearings

Use a hand bearing compass (minimally affected by ship’s magnetism) as a reference.

Steps:

  1. Anchor or stop the boat
  2. Send a person ashore with the hand bearing compass
  3. Both take bearings to each other simultaneously
  4. The two bearings should differ by exactly 180°
  5. Any difference is deviation in the steering compass

Deviation Table

After measuring on multiple headings, record the results:

Ship’s Head (KK)Deviation
000° (N)—
045° (NE)—
090° (E)—
135° (SE)—
180° (S)—
225° (SW)—
270° (W)—
315° (NW)—

Interpolate between measured headings for any course in between.


Tips

  • Deviation changes if you add/move metal objects, electronics, or speakers near the compass
  • Remeasure after any significant modification to the boat’s equipment
  • A deviation of less than ±3° is generally acceptable for coastal sailing
  • Steer clear of metal objects and active electronics when taking compass readings

Quick Reference

GoalFormulaRule
KK → KMKM = KK + devGo UP, ADD
KM → KRKR = KM + varGo UP, ADD
KR → KVKV = KR + snosStarboard: −, Port: +
KV → KRKR = KV − snosGo DOWN, SUBTRACT
KR → KMKM = KR − varGo DOWN, SUBTRACT
KM → KKKK = KM − devGo DOWN, SUBTRACT
Compass → Chart bearingTB = KB + dev + varGo UP, ADD
Plot position lineReciprocal = TB ± 180°Draw from landmark
DistanceD = V × t(min) / 60—
Timet = D × 60 / V—

Sailing Logbook

Totals:

  • As skipper: ~848NM
  • Totals: ~948NM

Season 2026

  • Distance: TBA NM

Season 2025

  • Distance: ~105 NM

Season 2023

  • Distance: 265 NM
  • Maximum speed: 5.9 kts
  • Average speed: 3.9 kts

Season 2022

  • Distance: 208 NM
  • Maximum speed: 7.6 kts
  • Average speed: 3.1 kts

Season 2021

  • Distance: 183 NM
  • Maximum speed: 9.9 kts
  • Average speed: 4.3 kts

Season 2020

  • Distance: 87 NM
  • Maximum speed: 7.1 kts
  • Average speed: 3.1 kts

Season 2019

  • Distance 100NM [^1]

[^1] 100NM sailed in baltic sea as training

2026

Marina Sibenik 30.05.2026 - 06.06.2026

Trip summary:

  • Distance: TBA NM
  • Maximum speed: - kts
  • Average speed: - kts
  • Boat: Dufour 460 GL | Almar
    • Year: 2017
    • Drought: 2.20 m
    • Length: 14.15 m
    • Beam: 4.50 m
    • Engine: 75 hp (55,16 kW)
    • Fuel tank: 250 l
    • Water tank: 530 l
    • Classical mainsail

Map

route2026

Route stops and details:


Planning


Sailing Notes

Generic Check

  • Passport / ID card
  • Driving licence
  • Boat booking confirmation
  • Insurance
  • Crew list
  • Skipper licence
  • Boat contract / charter agreement
  • Cash
  • Payment cards
  • Nautical charts & manual navigation tools

Personal Checklist

Clothing

  • Hat/Cap
  • Sunglasses + strap
  • UV protective long-sleeve shirt (sun on open water is intense)
  • Short-sleeve shirts
  • Shorts
  • Long trousers (evenings, March-April)
  • Swimwear
  • Light fleece / hoodie (cool nights, early season)
  • Windbreaker / windstopper
  • Light rain jacket (squalls occur even in summer)
  • Shoes for land + sandals / flip flops
  • Light-soled shoes or deck boots for the boat
  • Reef shoes / water shoes (rocky Med beaches)
  • Sailing gloves

Gear

  • Travel pillow
  • Headlight (with red light)
  • Snorkeling mask + fins
  • Binoculars
  • Powerbank
  • Hammock

Hygiene & Health

  • Towel (quick-dry)
  • Sunscreen SPF 50+
  • After-sun lotion
  • Lip balm with SPF
  • Insect repellent (mosquitoes are common in marinas at dusk)
  • Rehydration salts / electrolytes (heat exhaustion risk)
  • Antihistamines (fenistil — jellyfish stings, insect bites)
  • Eye drops (sun and salt air)
  • Soap / shower gel
  • Medicines
  • Nausea pills (ginger sweets / lollipops)

Entertainment & Extras

  • Musical instruments / songbooks
  • Books

Boat Supplies

Kitchen & Cooking

  • Aeropress / moka pot / French press
  • Thermos flask

Tools & Repair

  • Rope / string (1m short + longer line)
  • Matches x3 + lighter + candles
  • Needle & thread + plasters
  • Duck tape
  • Carpet tape
  • Screws
  • Phillips screwdriver
  • Soldering gun

Pre-departure Checklist

  • Knots practice: cleat hitch, clove hitch, figure-eight, bollard tie, bowline
  • Boat orientation:
    • Lines in lockers — which does what
    • Mainsail halyard
    • Furling jib
    • Main sheet & jib sheet
    • Winches
    • Anchor & controller
  • Autopilot is used only on engine or when waves are calm and the boat is well balanced from a sail trim perspective
  • Prepare helm for quick deployment
  • Always keep everything secured — the sea is not always calm and everything follows gravity
  • Install safety lines and safety net on railing (when Amelia is on board)

Rules

  1. The designated skipper is always right — he is currently responsible for safety and navigation
  2. Skipper is responsible for safety — follow instructions immediately, without discussion
  3. Do not stand on cabin steps

Handover inspection

  • Rudder
  • Boom joint
  • Furling jib lift
  • Aft leech tensioner
  • Navigational lights (port, starboard, masthead, stern, anchor light)
  • Navigational ball and triangle (day shapes)
  • Anchor (condition, chain, windlass operation)
  • Autopilot (function, drive unit)
  • Speedometer / log
  • GPS / chartplotter
  • Fuses (location, spares)
  • Tools (location, contents)
  • Engine room (belts, hoses, bilge, seacocks)
  • Oil level — engine (dipstick fully inserted to read) and saildrive (dipstick rested on thread, do not screw in)
  • Toilets & black water tank (check integrity with a bit of saponate — look for foam or food coloring at through-hulls)

Adriatic Winds

WindDirectionSeasonTypical ForceCharacterHazards & Notes
BuraNE – ENEOct – Apr; possible year-round6 – 10+ (gusts >60 kn)Cold, dry, katabatic; descends from Dinaric Alps; extremely gustyMost dangerous Adriatic wind; strongest in Kvarner, Velebit channel, and Trieste; can appear with little warning; short steep seas
JugoSE – SAutumn, winter, spring5 – 8Warm, humid, persistent; builds slowly over 1–3 days; long-period swellPoor visibility, rain; seas become very rough and confused; subsides slowly; watch for prolonged forecasts
MaestralNW – WNWJun – Sep (afternoons)3 – 5Thermal sea breeze; develops late morning, peaks early afternoon, dies at sunsetReliable summer sailing wind; pleasant conditions; no significant swell; can build to F5–6 on exposed coasts
TramontanaN – NNEAutumn, spring4 – 7Cold, dry, continental; steadier and less gusty than BuraBrings clear skies; chilly; can be confused with Bura onset
NeveraNW – N (squall)Jun – SepSquall: 7 – 10+Sudden violent thundersquall; forms rapidly over mountains in summer heatMost dangerous in summer; little warning; waterspouts possible; seek shelter immediately when cumulonimbus build inland
LebićSW – WSWAutumn, winter4 – 7Warm, humid; brings swell from the open MediterraneanRain and poor visibility; less common than Jugo; can combine with Jugo to create confused cross-seas
OstroS – SSEVariable3 – 6Warm, humid southerly; precursor or variant of JugoBrings mist and cloud; often transitions into full Jugo
BurinNE – variable (offshore)Summer nights1 – 3Light nighttime land breeze; calm, dry; fades at sunriseComplement to daytime sea breeze; useful for motoring out of anchorage at night
PulenatWVariable3 – 5Moderate westerly; relatively rare in central AdriaticCan bring short choppy seas on exposed passages
LevanatE – ENEVariable3 – 5Easterly; uncommon in central and southern AdriaticMay bring spray and choppy conditions; more frequent in northern Adriatic
GregoGrecoNE (between Tramontana and Levanat)Autumn, winter4 – 7Cold and dry; similar to Tramontana; can be gusty near capes
Vardar—NE – EWinter5 – 8Cold, dry katabatic wind funnelled through river valleys in Greece/N Macedonia; reaches southern Adriatic

2025

Marina Preveza 14.06.2025 - 28.06.2025

Trip summary:

  • Distance: ~105 NM
  • Maximum speed: - kts
  • Average speed: - kts
  • Boat: Bavaria Cruiser 41 | Anemoessa
    • Year: 2016
    • Drought: 1.85 m
    • Length: 13.10 m
    • Beam: 3.99 m
    • Engine: 55 hp
    • Fuel tank: 210 l
    • Water tank: 360 l
    • Furling mainsail

Map

route2025

Route stops and details:


2023

Marina Veruda 29.04.2023 - 06.05.2023

Trip summary:

  • Distance: 215 NM
  • Maximum speed: 6.08 kts
  • Average speed: 3.88 kts
  • Boat: Elan 40.1 Impression | Estela
    • Year: 2020
    • Drought: 1.80 m
    • Length: 11.83 m
    • Beam: 3.91 m
    • Engine: 40 hp (29.8 kW)
    • Fuel tank: 146 l
    • Water tank: 400 l
    • Classic mainsail

Map

route2023

Route stops and details:

Marina Punat 16.09.2023 - 23.09.2023

Trip summary:

  • Distance: ~50 NM
  • Maximum speed: - kts
  • Average speed: - kts
  • Boat: Dufour 430 | Catacea
    • Year: 2020
    • Drought: 2.10 m
    • Length: 13.23 m
    • Beam: 4.30 m
    • Engine: 75 hp (55.9 kW)
    • Fuel tank: 250 l
    • Water tank: 530 l
    • Classic mainsail

Map

route2023

Route stops and details:


2022


Marina Pomer 07.05.2022 - 14.05.2022

Trip summary:

  • Distance: 208 NM
  • Maximum speed: 7.5 kts
  • Average speed: 3.3 kts
  • Boat: Beneteau Oceanis 35 | DalMar
    • Year: 2016
    • Drought: 1.85m
    • Length: 9.99 m
    • Beam: 3.7 m
    • Engine: 29 hp (21.6 kW)
    • Fuel tank: 130 l
    • Water tank: 330 l
    • Classic mainsail

Map

route2022

Route stops and details:


2021


Marina Frapa Rogoznica 29.05.2021 - 05.06.2021

Trip summary:

  • Distance: 183 NM
  • Maximum speed: 9.9 kts
  • Average speed: 4.3 kts
  • Boat: Dufour 430 | Calando
    • Year: 2020
    • Drought: 2.10m
    • Length: 13.24 m
    • Beam: 4.3 m
    • Engine: 60 hp (44.7 kW)
    • Fuel tank: 200 l
    • Water tank: 380 l
    • Rolling mainsail

Map

route2021

Route stops and details:


2020


Vodice 05.09.2020 - 12.09.2020

Trip summary:

  • Distance: 87 NM
  • Maximum speed: 7.1 kts
  • Average speed: 3.1 kts
  • Boat: Jeanneau Sun Odyssey 30 | Espresso 1
    • Year: 2009
    • Drought: 1.95m
    • Length: 8.99 m
    • Beam: 3.18 m
    • Engine: 21 hp (15.7 kW)
    • Fuel tank: 50 l
    • Water tank: 160 l
    • Rolling mainsail

Route details


Sailing Galleries

Veruda - Dragove - Rovijn - Veruda 2023

Pomer - Veli Rat - Pomer 2022

Frapa - Vis - Frapa 2021

Vodice - Zut - Vodice 2020

Filesystem

XFS

XFS logging

Overview

XFS uses Write-Ahead Logging (WAL) to guarantee filesystem metadata consistency. Every metadata change is recorded in the log before being applied to the on-disk structures. If a crash occurs, the log is replayed to bring the filesystem back to a consistent state.

The logging subsystem has two major layers that work together:

  1. The circular on-disk log — a fixed-size ring buffer of 512-byte blocks stored in a dedicated log device (or the end of the data device).
  2. The Committed Item List (CIL) / Delayed Logging layer — an in-memory aggregation layer that batches and de-duplicates log writes before flushing to the on-disk log.

Key Data Structures

struct xlog — The Log Manager

Defined in xfs_log_priv.h, this is the central control structure for the entire logging subsystem.

struct xlog {
    struct xfs_mount        *l_mp;           // owning filesystem mount
    struct xfs_ail          *l_ailp;         // Active Item List
    struct xfs_cil          *l_cilp;         // Committed Item List (delayed logging)
    struct xlog_grant_head   l_reserve_head; // logical reservation accounting
    struct xlog_grant_head   l_write_head;   // physical space accounting
    atomic64_t               l_tail_lsn;     // LSN of oldest unpersisted transaction
    struct xlog_in_core     *l_iclog;        // head of the iclog ring
    spinlock_t               l_icloglock;    // protects the iclog state machine
};

The two grant_head fields are the heart of log space management and are explained in detail in Log Space Accounting.


struct xlog_in_core — The In-Core Log Buffer (iclog)

Each iclog is a chunk of memory that absorbs formatted log records before they are written to disk. They form a circular ring buffer of typically 4 to 8 buffers, each up to 256 KB.

struct xlog_in_core {
    enum xlog_iclog_state  ic_state;       // state machine position
    atomic_t               ic_refcnt;      // reference count
    wait_queue_head_t      ic_force_wait;  // waiters on forced flush
    struct xlog_in_core   *ic_next;        // next in ring
    u32                    ic_offset;      // current write cursor
    u32                    ic_size;        // total buffer size
    void                  *ic_datap;       // pointer to buffer data
    struct list_head       ic_callbacks;   // CIL checkpoint callbacks
};

Iclog state machine (xfs_log.c):

ACTIVE → WANT_SYNC → SYNCING → DONE_SYNC → CALLBACK → DIRTY → ACTIVE
StateMeaning
ACTIVEAccepting new log records
WANT_SYNCFull or flushed, waiting for all writers to finish
SYNCINGI/O submitted to disk
DONE_SYNCI/O complete, callbacks pending
CALLBACKRunning CIL checkpoint callbacks (AIL insertion)
DIRTYBuffer spent; being recycled back to ACTIVE

struct xfs_cil and struct xfs_cil_ctx — The Delayed Logging Layer

xfs_cil_ctx (xfs_log_priv.h) is the container for a single CIL checkpoint — a batch of log items accumulated since the last checkpoint flush.

struct xfs_cil_ctx {
    xfs_csn_t          sequence;    // monotonically increasing checkpoint number
    xfs_lsn_t          start_lsn;  // LSN of first log record
    xfs_lsn_t          commit_lsn; // LSN of commit record
    struct list_head   lv_chain;   // chain of formatted shadow buffers
    atomic_t           space_used; // bytes accumulated so far
    struct xlog_ticket *ticket;    // log reservation ticket
};

xfs_cil (xfs_log_priv.h) owns the current context and drives the push:

struct xfs_cil {
    struct xlog           *xc_log;
    struct rw_semaphore    xc_ctx_lock;   // write-locked during push only
    struct xfs_cil_ctx    *xc_ctx;        // current live context
    spinlock_t             xc_push_lock;  // protects ordering list
    wait_queue_head_t      xc_commit_wait;
    void __percpu         *xc_pcp;        // per-CPU item lists
};

struct xfs_log_vec — The Shadow Buffer

When a transaction commits, each log item’s in-memory state is formatted into a shadow buffer (a log_vec) and decoupled from the live object. This is the central innovation of delayed logging.

struct xfs_log_vec {
    struct list_head      lv_list;     // CIL chain
    uint32_t              lv_order_id; // intra-checkpoint ordering
    int                   lv_niovecs;  // number of iovecs
    struct xfs_log_iovec *lv_iovecp;  // formatted region descriptors
    struct xfs_log_item  *lv_item;    // back-pointer to log item
    char                 *lv_buf;     // shadow buffer memory
    int                   lv_bytes;   // bytes used
};

After formatting, the log item is unlocked immediately. The shadow buffer holds all data required for the eventual log write.


Log Space Accounting: The Dual-Grant-Head Model

XFS tracks log space with two independent accounting heads, both defined as xlog_grant_head (xfs_log_priv.h):

HeadWhat it tracksCan overcommit?
l_reserve_headLogical reservation (space promised to transactions)Yes — future commits block, but existing ones proceed
l_write_headPhysical bytes actually written to the logNo — hard limit, never advances past the tail

Why two heads? Rolling transactions (e.g., directory operations that span many buffer modifications) need to reserve space upfront without exhausting the physical log. The reserve head allows logical overcommitment so a rolling transaction can keep rolling, while the write head enforces the actual circular buffer boundary.

CIL space limits (xfs_log_priv.h):

// Background push triggered at ~12.5% of total log size
#define XLOG_CIL_SPACE_LIMIT(log)  min_t(int, (log)->l_logsize >> 3, ...)

// Transaction commits throttled (sleeping wait) at 25% of log size
#define XLOG_CIL_BLOCKING_SPACE_LIMIT(log)  (XLOG_CIL_SPACE_LIMIT(log) * 2)

LSN Encoding

Log Sequence Numbers are 64-bit values:

LSN = (cycle << 32) | block_offset
  • cycle: how many times the log has wrapped around.
  • block_offset: 512-byte block offset within the log.

Macros CYCLE_LSN() and BLOCK_LSN() extract these fields throughout the code.


Transaction Lifecycle with Delayed Logging

This is the full path from a filesystem operation to a durable on-disk record.

Phase 1: Transaction Allocation and Reservation

xfs_trans_alloc()
  → xlog_ticket_alloc()
  → xlog_grant_head_check()   ← may sleep here if log is full

A ticket (reservation) is allocated holding the worst-case byte count for this transaction type. The reservation is calculated at mount time from geometric properties (tree depth, block size, etc.) and accounts for recursive modifications.

Phase 2: Item Modification

The caller modifies in-memory metadata (inode, buffer, dquot). Items are logged via xfs_trans_log_inode(), xfs_trans_log_buf(), etc., which mark items dirty on the transaction.

Phase 3: Transaction Commit → CIL Insertion

xfs_trans_commit() → xlog_cil_commit():

  1. Shadow buffer allocation (xlog_cil_alloc_shadow_bufs()): outside xc_ctx_lock, allocate memory sized for each item’s formatted representation.
  2. Lock acquisition: acquire xc_ctx_lock as a reader (allows concurrent commits).
  3. Format into shadow buffers: call each item’s iop_format() callback, writing item state into the shadow buffer.
  4. Pin on first insertion: if the item has not been in the CIL before, call iop_pin(). Re-logging (subsequent modifications before a checkpoint) does not add additional pins — the existing pin is reused and the old shadow buffer is discarded.
  5. Add to per-CPU list: attach the log_vec to the per-CPU CIL pending list.
  6. Unlock item: the live object is immediately available for further modification.
  7. Release xc_ctx_lock.

Relogging

Relogging is the critical property that prevents log tail pinning. Each re-commit of an item supersedes the previous version. Only the latest aggregate snapshot is written to the on-disk log (xfs-delayed-logging-design.rst):

Transaction  What is logged    LSN
    A             A             X
    B            A+B           X+n        ← A's slot at X is now stale
    C           A+B+C          X+n+m

Phase 4: CIL Background Push

The CIL push worker (xlog_cil_push_work(), xfs_log_cil.c) is triggered when:

  • CIL space exceeds XLOG_CIL_SPACE_LIMIT (background push), or
  • A caller explicitly issues xfs_log_force() or fsync().

Push sequence:

1. Acquire xc_ctx_lock as WRITER  (excludes all new commits)
2. Swap context: xc_ctx points to a new empty context
3. Release xc_ctx_lock            ← concurrent commits resume on new context
4. Sort and aggregate log_vecs from the old context
5. Write start record and all item vectors to iclogs
6. Order commit record among concurrent checkpoints (xlog_cil_order_write)
7. Write commit record to iclog; submit iclog for I/O
8. On I/O completion: insert items into AIL, call iop_committed, unpin items

Step 6 — checkpoint ordering — ensures commit records appear in the log in strictly ascending checkpoint sequence order, regardless of when individual iclogs complete I/O.

Phase 5: Iclog I/O and AIL Insertion

When an iclog transitions from SYNCING to DONE_SYNC:

  1. xlog_state_iodone_process_iclog() runs callbacks.
  2. xlog_cil_process_committed() inserts log items into the Active Item List (AIL) at their commit LSN.
  3. Items are unpinned (iop_unpin()), making them eligible for writeback.

Phase 6: AIL Writeback and Log Tail Advancement

Once items are inserted into the AIL the xfsaild kernel thread takes over. The full mechanism is described in The Active Item List and Log Tail Pushing.

Log space is only reclaimed when the tail advances. This makes the log a true circular buffer.


Simple vs Rolling Transactions

XFS transactions fall into two categories based on how they interact with the log grant space reservation system.

Simple Transactions

A simple transaction reserves log space once, performs its work, and commits. The entire lifecycle fits within a single log reservation unit.

The reservation is computed as:

need_bytes = tr_logres × 1

At allocation (xfs_trans_alloc), xfs_log_reserve checks against l_reserve_head and advances both grant heads by need_bytes:

xlog_grant_add_space(&log->l_reserve_head, need_bytes);
xlog_grant_add_space(&log->l_write_head, need_bytes);

At commit, xfs_log_ticket_ungrant returns unused space from both heads:

bytes = t_curr_res;                /* unused portion */
xlog_grant_sub_space(&log->l_reserve_head, bytes);
xlog_grant_sub_space(&log->l_write_head, bytes);

The transaction ticket has t_cnt = 0 (no refills) and the XFS_TRANS_PERM_LOG_RES flag is not set.

Examples of simple transactions:

TransactionOperation
tr_fsynctsTimestamp update (utimensat, xfs_vn_update_time)
tr_sbSuperblock counter update
tr_swriteSynchronous inode write

Simple transactions never call xfs_trans_roll and never need xfs_log_regrant.


Rolling (Permanent) Transactions

A rolling transaction is allocated with XFS_TRANS_PERM_LOG_RES, signaling that it may need multiple log reservation refills. The initial reservation is:

need_bytes = tr_logres × tr_logcount

where tr_logcount represents the maximum number of “units” the transaction can use before needing to request more space from the grant heads.

The first unit goes to t_curr_res (the working reservation); the remaining tr_logcount - 1 units are held in t_cnt as pre-paid refills.

When a rolling transaction exhausts its current unit, it calls xfs_trans_roll() (xfs_trans.c), which performs a three-step handover:

  1. Duplicate — xfs_trans_dup() creates a new transaction structure, copying the ticket (with a reference count bump), transferring remaining block reservations and deferred operations to the new transaction. The old transaction’s XFS_TRANS_PERM_LOG_RES flag propagates to the new one.

  2. Commit the old transaction — __xfs_trans_commit(tp, true) commits with regrant = true. This tells the commit path to call xfs_log_ticket_regrant() instead of xfs_log_ticket_ungrant(). The difference is critical: regrant returns the unused portion of the current unit but retains the ticket for the next unit.

  3. Regrant log space — xfs_log_regrant() acquires the next unit of log space for the new transaction.

/* xfs_trans_roll() — simplified */
*tpp = xfs_trans_dup(tp);           /* step 1 */
error = __xfs_trans_commit(tp, true); /* step 2: regrant=true */
error = xfs_log_regrant(mp, (*tpp)->t_ticket);  /* step 3 */

The Regrant Path

xfs_log_ticket_regrant() (xfs_log.c) handles the grant head accounting at roll time:

void xfs_log_ticket_regrant(struct xlog *log, struct xlog_ticket *ticket)
{
    if (ticket->t_cnt > 0)
        ticket->t_cnt--;

    /* Return unused portion of current unit */
    xlog_grant_sub_space(&log->l_reserve_head, ticket->t_curr_res);
    xlog_grant_sub_space(&log->l_write_head, ticket->t_curr_res);
    ticket->t_curr_res = ticket->t_unit_res;  /* reset for next unit */

    /* If pre-paid refills remain, no need to acquire more space */
    if (!ticket->t_cnt) {
        /* Out of pre-paid units — must re-reserve on reserve head */
        xlog_grant_add_space(&log->l_reserve_head, ticket->t_unit_res);
    }
}

If pre-paid refills remain (t_cnt > 0), the refill is free — the space was already reserved at allocation time. When pre-paid refills run out (t_cnt == 0), xfs_log_regrant() must acquire new space from l_write_head via xlog_grant_head_check():

/* xfs_log_regrant() */
tic->t_tid++;
tic->t_curr_res = tic->t_unit_res;
if (tic->t_cnt > 0)
    return 0;      /* pre-paid refill — no grant check needed */

/* Out of pre-paid units: check the WRITE head */
error = xlog_grant_head_check(log, &log->l_write_head, tic, &need_bytes);

Note the critical difference in which grant head is checked:

PathFunctionGrant head checked
Initial reservationxfs_log_reserve()l_reserve_head (allows overcommit)
Regrant (out of units)xfs_log_regrant()l_write_head (hard physical limit)

The reserve head allows logical overcommitment — it promises space that may not yet be physically available. The write head enforces the actual circular buffer boundary. A rolling transaction that exhausts its pre-paid units and needs more space contends directly on the write head, which cannot advance past the log tail.

Why Rolling Transactions Exist

Operations like unlink, create, rename, and truncate can generate an unbounded number of log items through deferred operations. Deleting a file with 10,000 extents generates 10,000 EFI/EFD intent pairs. A single non-rolling transaction would need to reserve enough log space for all of them upfront — potentially exceeding the entire log.

Rolling transactions break this into bounded units. Each roll commits the work done so far (allowing the CIL to absorb it), releases the unused reservation, and obtains a fresh unit for the next batch. The xfs_defer_finish() mechanism drives this: it processes deferred ops in batches, rolling the transaction between batches.

xfs_trans_commit()
  → xfs_defer_finish_noroll()   ← processes deferred ops
      → xfs_defer_create_intents()  ← log intent items
      → xfs_trans_roll()            ← commit current, get fresh unit
      → xfs_defer_trans_roll()      ← re-join items to new transaction
      → xfs_defer_finish_one()      ← execute the deferred operation
      → xfs_defer_create_done()     ← log done item
      (loop until all deferred ops complete)

Rolling Transaction Types

All rolling transactions in XFS set XFS_TRANS_PERM_LOG_RES in tr_logflags:

Transactiontr_logcountOperation
tr_createper-mount computedFile/node creation
tr_removeper-mount computedUnlink
tr_renameper-mount computedRename
tr_linkper-mount computedHard link
tr_mkdirper-mount computedDirectory creation
tr_symlinkper-mount computedSymlink creation
tr_write2 (8 with reflink)Buffered write allocation
tr_itruncate2 (8 with reflink)Truncate
tr_ifree2Inode inactivation
tr_addafork2Add attribute fork
tr_create_tmpfile2O_TMPFILE creation
tr_growdata2Grow data section
tr_attrinval1Attribute invalidation

The per-mount computed counts (from functions like xfs_icreate_log_count()) factor in attribute operations that may accompany the primary operation, including potential B-tree splits at the filesystem’s actual tree depth. For reflink-enabled filesystems, tr_write and tr_itruncate increase from 2 to 8 units to accommodate CoW extent remapping and refcount B-tree updates.


Grant Head Interaction Summary

Simple transactions interact with the grant system once — a brief reservation on l_reserve_head lasting microseconds between xfs_trans_alloc() and xfs_trans_commit(). Under concurrency, thousands of simple transactions contribute minimal sustained grant head pressure because their reservations are allocated and released so quickly.

Rolling transactions interact with the grant system multiple times. Each roll returns unused space then re-acquires a full unit. When pre-paid units run out, the regrant path checks l_write_head, which is the harder limit — it cannot advance past the log tail. When the AIL cannot drain (slow storage), the write head fills and rolling transactions stall on regrant, creating sustained waiters on l_write_head.

Both paths funnel through the same xlog_grant_head_check() function, which manages the FIFO waiters queue. This has implications for fairness: a waiter on either head’s queue affects all newcomers entering the same head’s check path. The interaction between large rolling-transaction reservations and many small simple-transaction reservations competing on the same grant head queue is a source of performance pathology under log pressure.


The Active Item List and Log Tail Pushing

The AIL is the bridge between the log and the on-disk metadata. Its sole purpose is to track every log item that has been committed to the log but not yet written to its final on-disk location, and to push those items to disk so the log tail can advance and log space can be reclaimed.

Data Structures

struct xfs_ail (xfs_trans_priv.h)

struct xfs_ail {
    struct xlog          *ail_log;            // log being managed
    struct task_struct   *ail_task;           // xfsaild kthread
    struct list_head      ail_head;           // LSN-ordered item list
    struct list_head      ail_cursors;        // active traversal cursors
    spinlock_t            ail_lock;           // protects all AIL state
    xfs_lsn_t             ail_last_pushed_lsn;// LSN of last successfully pushed item
    xfs_lsn_t             ail_head_lsn;      // log head LSN at AIL init
    int                   ail_log_flush;      // counter: force CIL push when set
    unsigned long         ail_opstate;        // XFS_AIL_OPSTATE_PUSH_ALL flag
    struct list_head      ail_buf_list;       // buffers queued for delwri submission
    wait_queue_head_t     ail_empty;          // waiters for AIL to drain completely
    xfs_lsn_t             ail_target;        // LSN we are currently pushing toward
};

The list at ail_head is kept in strict ascending LSN order. The item at the front (minimum LSN) defines the log tail: that is the oldest record in the log that has not yet been written to disk.

struct xfs_ail_cursor (xfs_trans_priv.h)

struct xfs_ail_cursor {
    struct list_head      list;   // registered in ailp->ail_cursors
    struct xfs_log_item  *item;   // current position (low bit = invalidated)
};

Cursors allow xfsaild to walk the AIL safely even when items are deleted concurrently. When an item is removed, every cursor pointing at it has its low pointer bit set. The next call to xfs_trans_ail_cursor_next() detects this and restarts the traversal from the new minimum.


The xfsaild Daemon Main Loop (xfs_trans_ail.c:653)

xfsaild is a single per-filesystem kthread. Its loop has three states:

┌──────────────────────────────────────────────────────┐
│ Set TASK_KILLABLE or TASK_INTERRUPTIBLE               │
│   (KILLABLE if tout ≤ 20ms for fast wakeup)          │
├──────────────────────────────────────────────────────┤
│ Check kthread_should_stop() → drain ail_buf_list      │
│   and exit on shutdown                                │
├──────────────────────────────────────────────────────┤
│ If AIL empty AND ail_buf_list empty → schedule()     │
│   (full idle: no timeout, wait for wakeup)            │
├──────────────────────────────────────────────────────┤
│ If tout > 0 → msleep(tout)                           │
├──────────────────────────────────────────────────────┤
│ tout = xfsaild_push(ailp)  ← core work               │
└──────────────────────────────────────────────────────┘

The thread is woken by:

  • xfs_ail_push() — called from xlog_assign_tail_lsn() when the log approaches full.
  • xfs_ail_push_all() — called by umount and log quiesce to drain the AIL completely.
  • xfs_trans_ail_update_bulk() — any new insertion into the AIL.

Push Target Calculation (xfs_trans_ail.c:405)

Before scanning items, xfsaild_push() calls xfs_ail_calc_push_target() to decide how far to push. The logic in order of priority:

  1. Push-all flag set (XFS_AIL_OPSTATE_PUSH_ALL) or ail_empty has waiters: return max_lsn (the current log head). Push everything.

  2. Log already has ≥ 25% free space:

    free_bytes = l_logsize − (head_lsn − min_lsn)
    if free_bytes ≥ l_logsize / 4 → keep current ail_target
    

    No pushing needed; keep the existing target.

  3. Log has < 25% free space: advance the target by 25% of the log size from the current tail:

    target_block = BLOCK_LSN(min_lsn) + (l_logBBsize >> 2);
    // wrap cycle if needed
    target_lsn   = xlog_assign_lsn(target_cycle, target_block);
    

    The target is clamped to max_lsn and never lowered below the existing ail_target.

Design intent: one push round reclaims exactly 25% of the log, ensuring a predictable amount of free space without over-flushing.


The xfsaild_push() Loop (xfs_trans_ail.c:535)

This is the core of the push. It runs under ail_lock for the traversal, briefly dropping it during I/O operations.

Step 1: CIL Pre-flush Optimization

if (ailp->ail_log_flush && ailp->ail_last_pushed_lsn == 0 &&
    (!list_empty_careful(&ailp->ail_buf_list) || xfs_ail_min_lsn(ailp))) {
    ailp->ail_log_flush = 0;
    xlog_cil_flush(ailp->ail_log);
}

When the AIL has items but the push cursor is stuck at the beginning (ail_last_pushed_lsn == 0), it means items are pinned by in-flight CIL transactions. Rather than spinning on pinned items, xfsaild forces a synchronous CIL flush. This breaks the potential circular wait:

CIL holds pins → AIL cannot advance → log fills → CIL cannot commit → deadlock

Step 2: Cursor Initialization and Target Update

WRITE_ONCE(ailp->ail_target, xfs_ail_calc_push_target(ailp));
lip = xfs_trans_ail_cursor_first(ailp, &cur, ailp->ail_last_pushed_lsn);

The cursor starts from ail_last_pushed_lsn so that a push that hit the item limit in one round can continue from where it left off in the next.

Step 3: Item Traversal

while (XFS_LSN_CMP(lip->li_lsn, ailp->ail_target) <= 0) {
    if (test_bit(XFS_LI_FLUSHING, &lip->li_flags))
        goto next_item;             // skip: already in-flight

    xfsaild_process_logitem(ailp, lip, &stuck, &flushing);
    count++;

    if (stuck > 100)
        break;                      // backoff: too many blocked items
    if (lip->li_lsn != lsn && count > 1000)
        break;                      // per-LSN limit: avoid infinite loop
}

Two hard limits prevent the push loop from monopolizing the CPU:

  • stuck > 100: if more than 100 consecutive items are pinned or locked, abort and sleep. Continuing would just burn CPU with no progress.
  • count > 1000 at a new LSN: prevents unbounded iteration when many items share the same commit LSN.

Step 4: Async Buffer Submission

if (xfs_buf_delwri_submit_nowait(&ailp->ail_buf_list))
    ailp->ail_log_flush++;

All buffers queued during the traversal are submitted in a single batched write. submit_nowait returns non-zero if the submission was not possible (e.g. I/O error or congestion), which sets ail_log_flush to trigger a CIL flush on the next round.

Step 5: Timeout Selection

The return value controls how long xfsaild sleeps before the next round:

ConditiontoutMeaning
Reached target, or AIL empty50 msWait for in-flight I/O to complete; reset cursor to 0
>90% of items were stuck/flushing20 msBack off; next round may issue a log force; reset cursor to 0
More items remain below target0 msReturn immediately; continue from ail_last_pushed_lsn

The cursor reset (ail_last_pushed_lsn = 0) on the first two cases ensures the next wakeup re-evaluates the entire AIL from the minimum, picking up items that may have been unpinned during the sleep.


Per-Item Push: xfsaild_process_logitem() (xfs_trans_ail.c:468)

For each item in the traversal, xfsaild_push_item() dispatches to the item’s iop_push callback and interprets the return code:

Return codeMeaningAction
XFS_ITEM_SUCCESSQueued for I/OUpdate ail_last_pushed_lsn
XFS_ITEM_FLUSHINGAlready being writtenIncrement flushing; update ail_last_pushed_lsn
XFS_ITEM_PINNEDHeld by an in-memory transactionIncrement stuck; set ail_log_flush
XFS_ITEM_LOCKEDCould not acquire buffer lockIncrement stuck
XFS_ITEM_FAILEDPrevious I/O failedResubmit via xfsaild_resubmit_item()

Items with XFS_LI_FAILED set are handled by xfsaild_resubmit_item() which re-queues the backing buffer directly to ail_buf_list without calling iop_push again, allowing the I/O to be retried on the next submission round.


Inode Item Push: xfs_inode_item_push() (xfs_inode_item.c:739)

Inode items use cluster flushing to amortize I/O overhead. A cluster is a group of inodes that share a single filesystem buffer (typically a 4 KB or 16 KB block).

1. Check preconditions (return PINNED or FLUSHING if not ready):
   - inode stale (being freed)?    → PINNED
   - ipincount > 0?                → PINNED
   - cluster buffer pinned?        → PINNED
   - XFS_IFLUSHING flag set?       → FLUSHING
   - xfs_buf_trylock() fails?      → LOCKED

2. Release ail_lock              ← avoids holding spinlock during I/O
3. xfs_iflush_cluster(bp)       ← formats ALL inodes in the cluster into bp
4. xfs_buf_delwri_queue(bp, &ailp->ail_buf_list)
5. Reacquire ail_lock

xfs_iflush_cluster() walks all inodes mapped to the same buffer and formats each one’s in-memory xfs_dinode into the buffer in one pass. This means that when xfsaild pushes one inode item, it potentially writes dozens of inodes with a single I/O, which is critical for performance on inode-dense workloads.

The AIL lock is dropped during xfs_iflush_cluster(). Cursors handle any concurrent deletions that occur during this window.


Buffer Item Push: xfs_buf_item_push() (xfs_buf_item.c:565)

Buffer items (btree blocks, superblock, AGF/AGI headers, etc.) have a simpler push path:

1. xfs_buf_ispinned(bp)?     → PINNED   (transaction holds a log reference)
2. xfs_buf_trylock(bp) fails?→ LOCKED   (re-check pin after trylock failure)
3. Log a warning if XBF_WRITE_FAIL is set (previous write error)
4. xfs_buf_delwri_queue(bp, &ailp->ail_buf_list)
5. xfs_buf_unlock(bp)

Unlike inodes, each buffer item maps 1:1 to a buffer, so no clustering is needed. The trylock avoids blocking — if the buffer is locked by another writer, xfsaild moves on and returns to it on the next round.


Tail Advancement: __xfs_ail_assign_tail_lsn() (xfs_trans_ail.c:753)

When xfs_ail_delete() removes an item from the AIL, it calls xfs_ail_update_finish(), which calls __xfs_ail_assign_tail_lsn():

tail_lsn = __xfs_ail_min_lsn(ailp);   // LSN of first item in AIL
if (!tail_lsn)
    tail_lsn = ailp->ail_head_lsn;    // AIL empty: tail = current head

WRITE_ONCE(log->l_tail_space,
    xlog_lsn_sub(log, ailp->ail_head_lsn, tail_lsn));
atomic64_set(&log->l_tail_lsn, tail_lsn);

After updating the tail, xfs_ail_update_finish() calls xfs_log_space_wake(), which wakes all threads sleeping on l_reserve_head or l_write_head. This directly unblocks stalled transaction allocations.

The tail can only move forward. It is the minimum LSN of all items still in the AIL. The log space available to new transactions is:

available = l_logsize − (l_tail_space)
          = l_logsize − (head_lsn − tail_lsn)

Every item flushed to disk shrinks l_tail_space, freeing space for the next wave of transactions.


AIL Locking Summary

LockScopeNotes
ail_lock (spinlock)All AIL list/state accessDropped during I/O in inode push
xfs_buf.b_lockIndividual buffer stateAcquired via trylock only; never spins
i_pincount / b_pin_countPin reference countsAtomic; checked before attempting push

The key design rule: ail_lock is never held while waiting for I/O. It is dropped before xfs_iflush_cluster() and reacquired immediately after, with cursors protecting traversal safety across the gap.


On-Disk Log Format

Log Record Header (xfs_log_format.h)

struct xlog_rec_header {
    __be32  h_magicno;      // 0xFEEDbabe
    __be32  h_cycle;        // wrap count
    __be32  h_version;      // log version (1 or 2)
    __be32  h_len;          // data length in bytes
    __be64  h_lsn;          // this record's LSN
    __be64  h_tail_lsn;     // oldest uncommitted LSN at write time
    __le32  h_crc;          // CRC-32c of entire record
    __be32  h_num_logops;   // count of operations in this record
    __be32  h_cycle_data[]; // cycle number embedded in each 512-byte block
    uuid_t  h_fs_uuid;      // filesystem UUID
};

Operation Header

struct xlog_op_header {
    __be32  oh_tid;      // transaction ID (for grouping ops)
    __be32  oh_len;      // payload length
    __u8    oh_clientid; // XFS_TRANSACTION = 0x69
    __u8    oh_flags;    // START_TRANS | COMMIT_TRANS | CONTINUE_TRANS
};

Log Item Types

TypeValueDescription
XFS_LI_INODE0x123bInode core and data fork
XFS_LI_BUF0x123cRaw buffer (btree blocks, superblock, etc.)
XFS_LI_DQUOT0x123dQuota record
XFS_LI_EFI/EFD0x1236/7Extent free intent/done
XFS_LI_RUI/RUD0x123a/9Rmap update intent/done
XFS_LI_CUI/CUD0x123f/gRefcount update intent/done
XFS_LI_BUI/BUDintent pairsBMBT update intent/done
XFS_LI_ATTRI/ATTRDintent pairsXattr update intent/done

Intent/Done pairs implement a two-phase commit protocol for complex operations that span multiple sub-transactions (e.g., freeing extents requires updating the free space B-tree and the reverse-mapping B-tree). If the filesystem crashes between writing the Intent and the Done record, recovery re-executes the operation from the Intent.


B-tree Splits

A B-tree split is the most log-intensive operation in the XFS metadata path. A single record insertion can trigger a cascade of splits from leaf to root, each allocating a new block and logging multiple buffers. Because reservation sizes are calculated from the worst-case split depth, understanding splits is essential for understanding why XFS log reservations are as large as they are.


When a Split Occurs

XFS B-trees are full B+ trees: every block is kept as full as possible during insertion. When an insertion targets a block that is already at maximum capacity, the kernel first tries two cheaper alternatives before resorting to a split (xfs_btree_make_block_unfull(), xfs_btree.c):

  1. Left shift (xfs_btree_lshift()): move the leftmost record to the left sibling if it has space.
  2. Right shift (xfs_btree_rshift()): move the rightmost record to the right sibling if it has space.
  3. Split (xfs_btree_split()): only if both siblings are also full.

A split always produces exactly one new block at the current level and returns one new key/pointer pair to the caller, which must then insert that pair into the parent level — potentially triggering another split.


On-Disk Block Format

Every XFS B-tree block on disk begins with struct xfs_btree_block (libxfs/xfs_btree_format.h):

struct xfs_btree_block {
    __be32  bb_magic;    // per-btree magic (e.g. XFS_BNOBT_MAGIC)
    __be16  bb_level;    // 0 = leaf, 1+ = internal node
    __be16  bb_numrecs;  // number of records/keys currently stored
    union {
        struct xfs_btree_block_shdr s;  // AG-rooted trees (32-bit sibling ptrs)
        struct xfs_btree_block_lhdr l;  // inode-rooted trees (64-bit sibling ptrs)
    } bb_u;
};

Both header variants contain:

FieldPurpose
bb_leftsibBlock number of left sibling (or NULLAGBLOCK/NULLFSBLOCK)
bb_rightsibBlock number of right sibling
bb_blknoPhysical block address of this block
bb_lsnLSN of the last transaction that modified this block
bb_uuidFilesystem UUID (guards against cross-filesystem recovery)
bb_ownerAG number (AG-rooted) or inode number (inode-rooted)
bb_crcCRC-32c of the block (recalculated on recovery, not logged)

Following the header, a block contains either:

  • Leaf: a flat array of fixed-size records.
  • Internal node: an array of keys followed by an array of n+1 child pointers.

All integer fields are big-endian on disk.


The Split Mechanism: __xfs_btree_split()

The core implementation lives in __xfs_btree_split() (xfs_btree.c). For BMBT (block map B-tree) splits where no AGF lock is held, a worker-thread wrapper xfs_btree_split() offloads the call to avoid unbounded kernel stack growth during recursive allocation; all other tree types call __xfs_btree_split() directly.

Step 1: Allocate the New Right Block

xfs_btree_alloc_block(cur, &lptr, &rptr, stat)
xfs_btree_get_buf_block(cur, &rptr, &right, &rbp)
xfs_btree_init_block_cur(cur, rbp, level, 0)

Block allocation is type-specific:

B-treeSource of new blockSide effect logged
BNOBT / CNTBTxfs_alloc_get_freelist() — AG free listAGF header (XFS_AGF_FLFIRST, XFS_AGF_FLCOUNT)
INOBT / FINOBTxfs_alloc_vextent_near_bno() — AG free spaceAGF + AGI block counter
RMAPBTxfs_alloc_get_freelist() — AGFLAGF agf_rmap_blocks, space reservation
BMBTxfs_alloc_vextent_near_bno() — data AGInode fork block count

Every one of these block sources modifies an AG header (AGF or AGI), which is itself logged as a XFS_LI_BUF item. A split thus always generates at least two logged buffers before any tree data is touched.

Step 2: Divide Records Between Left and Right

lrecs = xfs_btree_get_numrecs(left)
rrecs = lrecs / 2
if (lrecs is odd && cursor position <= rrecs + 1)
    rrecs++          // tilt balance toward right when cursor is nearby
src_index = lrecs - rrecs + 1
xfs_btree_set_numrecs(left,  lrecs - rrecs)
xfs_btree_set_numrecs(right, rrecs)

The split point is chosen so that both blocks end up roughly half full. The odd- record tilt biases records toward the block the cursor is about to insert into, minimising the chance of an immediate follow-up split.

For leaf blocks: xfs_btree_copy_recs() copies the upper half of records into the right block.

For internal nodes: xfs_btree_copy_keys() and xfs_btree_copy_ptrs() copy the upper half of keys and their associated child pointers.

The key at src_index (the lowest key of the right block) is extracted and returned to the caller as the split key — the value that must be inserted into the parent level as the separator between left and right.

Step 3: Log Every Modified Buffer

This is the critical point at which the split becomes durable. The following log operations happen in order:

WhatFields loggedxfs_btree_log_block() flags
Right block — all header fieldsmagic, level, numrecs, both sibling ptrs, blkno, LSN, UUID, ownerXFS_BB_ALL_BITS (excludes bb_crc)
Right block — data (leaf)records 1..rrecsvia xfs_btree_log_recs()
Right block — data (node)keys 1..rrecs, ptrs 1..rrecsvia xfs_btree_log_keys() + xfs_btree_log_ptrs()
Left block — changed headerbb_numrecs, bb_rightsibXFS_BB_NUMRECS | XFS_BB_RIGHTSIB
Right-right sibling (if exists)bb_leftsibXFS_BB_LEFTSIB

xfs_btree_log_block() converts the field bitmask to a byte range and calls xfs_trans_log_buf(), which marks that range dirty in the transaction’s log vector. xfs_trans_buf_set_type() is called first to stamp the buffer as XFS_BLFT_BTREE_BUF, which recovery uses to distinguish B-tree blocks from other buffer types.

The CRC (bb_crc) is deliberately not logged. It is recalculated from the block contents during recovery using xfs_btree_reada_bufs(), ensuring the stored CRC always matches what is actually on disk after replay.

Step 4: Update Sibling Chain

Before logging, the sibling doubly-linked list is repaired:

Before split:
  [left] ↔ [right-right]

After split:
  [left] ↔ [right (new)] ↔ [right-right]

Three pointer writes are needed:

  • left->bb_rightsib = right (logged as part of left block header)
  • right->bb_leftsib = left (logged as part of right block header, XFS_BB_ALL_BITS)
  • right->bb_rightsib = right-right (logged as part of right block header)
  • right-right->bb_leftsib = right (logged separately: XFS_BB_LEFTSIB only)

The right-right block read uses xfs_btree_read_buf_block(), which may issue a synchronous read if the block is not already in the buffer cache. On cold-cache workloads this is a significant latency source.


Upward Propagation: Recursive Splits

After __xfs_btree_split() returns, the caller (xfs_btree_insrec()) must insert the split key and right-block pointer into the parent level. If the parent is also full, it too must split. xfs_btree_insert() drives this loop:

do {
    error = xfs_btree_insrec(cur, level, &nptr, &rec, &key, &ncur, &i);
    // nptr is non-null if a split occurred at this level
    level++;
} while (!xfs_btree_ptr_is_null(cur, &nptr));

Each iteration may allocate one block and log three to five buffers. The loop terminates only when a level has room for the new key/pointer without splitting.

Worst-case depth: on a filesystem with a large allocation group and all optional B-trees enabled, a fully-populated RMAPBT can reach five levels. A single extent allocation that triggers a split at every level logs five new blocks plus five parent block updates plus five AG header updates — thirty or more buffer log items for one allocation.

The per-transaction log reservation must cover this worst case upfront, which is why reservation sizes are computed from tree height and block size at mount time rather than at runtime.


Root Split: Growing the Tree

When the split reaches the root, there is no parent to absorb the new key. The tree must grow one level taller.

AG-Rooted Trees (BNOBT, CNTBT, INOBT, FINOBT, RMAPBT)

xfs_btree_new_root() (xfs_btree.c):

1. Allocate a new block  → becomes the new root
2. xfs_btree_set_root(cur, &nptr, +1)
     → update AG header (AGF or AGI) root pointer and level field
     → log AGF/AGI with XFS_AGF_ROOTS | XFS_AGF_LEVELS
3. Initialize new root block (level = old_height, numrecs = 2)
4. Log new root block:  XFS_BB_ALL_BITS
5. Copy lowest key of each child into new root keys
6. Log keys:  xfs_btree_log_keys(cur, nbp, 1, 2)
7. Write left-child and right-child pointers into new root
8. Log ptrs: xfs_btree_log_ptrs(cur, nbp, 1, 2)
9. Advance cursor: bc_nlevels++

The AG header (AGF or AGI) records the new root block number and the new tree height. On the next mount, XFS reads those fields to reconstruct the cursor starting position without scanning the tree.

Inode-Rooted Trees (BMBT)

xfs_btree_new_iroot() (xfs_btree.c) handles the bmap B-tree, where the root lives directly inside the inode fork rather than in a separate block:

1. Allocate a new block  → receives a copy of current inode-root contents
2. memcpy(new_block, inode_root_data)
   Fix bb_blkno in new block to match its physical address
3. Compress the inode fork to hold only the new root (one key + one pointer)
4. Log new child block:  XFS_BB_ALL_BITS + records/keys/ptrs
5. Log inode:  XFS_ILOG_CORE | xfs_ilog_fbroot(whichfork)
   (the inode fork data region is now a single-entry root node)
6. bc_nlevels++

The inode fork has a fixed size defined by its di_forkoff. Once the in-inode root cannot hold another key/pointer pair even after a split, the inode root gains another level outward, eventually consuming the entire fork and forcing a fork conversion.


Cursor Tracking Across a Split

The xfs_btree_cur maintains one xfs_btree_level entry per tree level, each holding a (buffer, position) pair:

struct xfs_btree_level {
    struct xfs_buf *bp;   // buffer holding the block at this level
    uint16_t        ptr;  // 1-based index of current key/record
};

After __xfs_btree_split() divides the block, the cursor position may have moved to the right block:

if (cur->bc_levels[level].ptr > lrecs + 1) {
    xfs_btree_setbuf(cur, level, rbp);           // switch to right block
    cur->bc_levels[level].ptr -= lrecs;          // adjust position
}

If there are more levels above, a second cursor is duplicated (xfs_btree_dup_cursor()). One cursor tracks the left child; the duplicated cursor’s parent-level pointer is incremented by one to track the right child. The insertion loop in xfs_btree_insert() manages which cursor to use at each level and deletes the spare when the split chain resolves.


What Gets Logged Per Split Level

Summing the buffer log items produced by one complete split at a single level:

BufferLog itemsCondition
AG header (AGF or AGI)XFS_LI_BUFAlways — block allocation modifies AG header
New right block headerXFS_LI_BUF (XFS_BB_ALL_BITS)Always
New right block data (recs/keys/ptrs)XFS_LI_BUFAlways
Left block headerXFS_LI_BUF (XFS_BB_NUMRECS | XFS_BB_RIGHTSIB)Always
Right-right sibling headerXFS_LI_BUF (XFS_BB_LEFTSIB)Only if right-right exists

That is four or five XFS_LI_BUF items per split level, plus the AG header items from block allocation, which may themselves modify additional blocks (e.g., the AGFL block list used by BNOBT/RMAPBT to store free blocks). In a five-level tree where every level splits, a single insertion can log twenty or more distinct buffers before the transaction commits.


Crash Recovery of a Split

Because every buffer modified during a split is logged before the transaction commits, crash recovery is straightforward:

  • Crash before commit: no log records for this transaction are durable. The pre-split blocks are unmodified. The newly allocated block may appear in the block allocation structures but will be reclaimed by the space recovery pass.
  • Crash after commit: all log records are durable. Recovery replays each buffer item in order, restoring every block to its post-split state. The B-tree is structurally consistent at the end of replay.

There are no intent records for splits. A split is not a deferred operation: it is atomic within the transaction that triggered it. Either all split records are committed together, or none of them are.


Crash Recovery

xfs_log_recover.c implements recovery in two passes.

Pass 1: Log Scanning

xlog_recover() walks the log from tail_lsn forward:

  • Reads each log record header.
  • Validates magic number and CRC.
  • Groups operation headers by transaction ID.
  • Builds an in-memory map of all transactions present in the log.

Pass 2: Replay

xlog_recover_commit_trans() replays each complete transaction:

  1. For each log item, call xlog_recover_commit_buffer/inode/dquot().
  2. Overwrite on-disk metadata with the logged versions.
  3. For intent items (EFI, RUI, etc.), reconstruct the pending operation and schedule it for Phase 2 completion via deferred operations.

xlog_recover_finish() processes all deferred operations, completing any partially-done multi-step operations (extent frees, rmap updates, etc.).

Invariant enforced by design: a CIL checkpoint must be smaller than half the total log size. This guarantees that at least one full checkpoint is always present in the log, making partial-write crashes safe.


Locking Hierarchy

Violating this order causes deadlock.

1. xfs_mount.m_sb_lock        (filesystem-wide, rarely held)
2. xfs_buf.b_lock             (individual buffer locks)
3. xfs_inode.i_lock           (inode lock)
4. xlog.l_icloglock           (spinlock, iclog state machine)
5. xfs_cil.xc_ctx_lock        (rwsem, CIL context switch)
6. xfs_cil.xc_push_lock       (spinlock, checkpoint ordering list)
7. xfs_ail.ail_lock           (spinlock, AIL list)

xc_ctx_lock is a sleeping rwsem, not a spinlock, specifically to avoid holding a spinlock during log I/O submission, which can sleep.


Performance Characteristics and Bottlenecks

Batching Efficiency (CIL)

The CIL’s primary value is write amplification reduction. Without delayed logging, each fsync or log force flushes all dirty items individually. With CIL, items modified 100 times between two checkpoints are written to disk once, at their final state. In workloads with heavy relogging (directory updates, quota updates), this can reduce log I/O by an order of magnitude.

Bottleneck: If the CIL is too small (< 8 MB on a busy filesystem), background pushes fire too frequently, destroying the batching benefit and driving up log I/O.


Per-CPU CIL Aggregation

Transaction commits add items to per-CPU pending lists (xc_pcp) to eliminate contention on a single lock. Space accounting uses per-CPU counters until the soft limit is approached, at which point it transitions to atomic operations and wakes the push worker.

Bottleneck: On workloads with very many small transactions (e.g., millions of small file creates), the per-CPU-to-atomic transition point creates a serialization spike. Threads pile up in xlog_cil_commit() contending on xc_ctx_lock write acquisition during the context switch.


Grant Head Waiters

When the log is full (write head nearly meets the tail), new transactions block in xlog_grant_head_wait() on a FIFO wait queue.

Bottleneck — log tail pinning: The tail can only advance when the AIL empties items. The AIL can only empty items when their buffers are written to disk. If the storage device is slow, the log fills up and all new transactions stall. This is the primary throughput bottleneck on write-heavy workloads on slow devices.

Bottleneck — reservation overestimation: Reservations are computed for the worst case (maximum B-tree depth). On a mostly-empty filesystem, actual usage is much less, but the reservation holds the full amount until released. This reduces parallelism on small logs.


Iclog Contention (l_icloglock)

Every thread completing a CIL commit must briefly hold l_icloglock to copy its log vectors into the current iclog and advance the write cursor. On many-core systems (32+ CPUs), this spinlock becomes a serialization point under high log bandwidth.

Bottleneck: Large CIL checkpoints writing megabytes of log data while holding l_icloglock for each 32 KB iclog block starve concurrent threads trying to start new transactions.


AIL Push Rate and Tail Stall

xfsaild is a single-threaded daemon. On systems with many concurrent metadata writers, it must push items fast enough to keep the tail advancing ahead of the write head.

Bottleneck — device throughput: If the block device cannot sustain the required writeback rate, xfsaild builds up a backlog, the AIL grows, the tail stalls, the write head catches the tail, and transaction allocation blocks — a global freeze. The only remedy is faster storage, a larger log, or reducing metadata write amplification.

Bottleneck — pinned items: Items held by long-running or stalled transactions cannot be pushed regardless of device speed. When xfsaild encounters more than 100 pinned items in a row it backs off (20–50 ms sleep). If the items remain pinned across many rounds, ail_log_flush accumulates and each new push round opens with a forced CIL flush to try to unpin them. A transaction that holds its locks too long effectively pins the log tail and starves all other writers.

Bottleneck — buffer lock contention: xfsaild uses trylock on all buffers and returns XFS_ITEM_LOCKED immediately if the lock is unavailable. Under heavy concurrent writeback, many buffers may be locked by page writeback or other kernel paths. Items returning LOCKED count against the stuck threshold (100 items), triggering the backoff before the target LSN is reached.

Bottleneck — single-threaded design: xfsaild processes the AIL serially. Each iteration submits buffers via xfs_buf_delwri_submit_nowait(), which is asynchronous, but the traversal itself is sequential. On workloads that produce millions of small dirty metadata items, the daemon can spend more time traversing the list than the device spends doing I/O. There is no parallelism within a single push round.

Bottleneck — cluster flush overhead: Inode pushes call xfs_iflush_cluster() which formats all inodes in a buffer cluster. While this amortizes I/O, it also means xfsaild drops and reacquires ail_lock for every inode buffer, and any cursor invalidation during that window forces a restart of the traversal from the AIL minimum. On a filesystem with millions of recently-modified inodes spread across many clusters, this restart overhead can significantly slow the effective push rate.

Observable symptoms of a stalled tail:

  • xfs_log_force latency increases (callers sleeping on l_write_head).
  • xfs_buf_delwri_submit_nowait returns non-zero repeatedly (sets ail_log_flush each time), causing redundant CIL flushes.
  • /proc/fs/xfs/stat counters xs_push_ail_pinned and xs_push_ail_locked grow faster than xs_push_ail_success.
  • xfsaild wakes with tout=20 continuously (>90% contention threshold crossed).

Tuning levers:

  • Larger log: more space between head and tail gives xfsaild more time before the write head catches the tail.
  • Dedicated log device (separate fast NVMe): isolates log writes from data writeback, reducing contention on the device queue.
  • vm.dirty_ratio / vm.dirty_background_ratio: reducing the dirty page ratio limits how many buffers can be in-flight at once, reducing lock contention seen by xfsaild.

Checkpoint Ordering Serialization

xlog_cil_order_write() (xfs_log_cil.c) ensures commit records are written in sequence order. When two concurrent checkpoints race, the higher-sequence one must wait for the lower-sequence one to establish its commit_lsn before writing its own commit record.

Bottleneck: Under extreme concurrency with many small checkpoints firing in rapid succession, checkpoint ordering serialization limits the rate at which new commit LSNs can be established, capping throughput in the log-write path.


Recovery Time

Recovery time is proportional to the amount of data between tail_lsn and head_lsn at the time of crash. A larger log retains more history, meaning more data to replay. On systems with very large logs (hundreds of GB) and high write rates before the crash, recovery can take minutes.


Bottlenecks and Write Amplification

Write amplification in XFS logging occurs at several independent layers. Each layer multiplies the number of actual device writes relative to the application-level operation that triggered them. Understanding which layer is responsible for observed I/O load is essential for diagnosis.


Write Amplification Taxonomy

Layer 1: Fundamental WAL Amplification

Every metadata modification is written twice: once sequentially to the log, and once in-place to the metadata location on disk. This is the irreducible cost of crash consistency via WAL. A single mkdir that modifies an inode, a directory block, and two AGF entries produces at minimum four log writes and four eventual on-disk writes — eight device writes for four logical changes.

Application write
  └─ metadata change
       ├─ → log write (sequential, via iclog)
       └─ → on-disk write (random, via AIL writeback)

The log write is sequential and cheap per-byte. The on-disk write is random and expensive per-operation. For metadata-heavy workloads on rotational storage the random on-disk writes dominate; on NVMe the log bandwidth is more often the limit.

Layer 2: Relogging Amplification (Pre-CIL)

Before delayed logging, every transaction that modified an already-logged item wrote the item to the log again in full, even if the change was a single byte. A hot inode touched by 1 000 transactions before being flushed to disk would appear 1 000 times in the log. Log space consumption was proportional to transaction count, not to the number of distinct objects.

CIL eliminates this at the log level: the item is formatted once per checkpoint regardless of how many transactions modified it within that checkpoint window. The reduction in log write amplification depends entirely on the relogging rate. On workloads with heavy relogging (directory entry updates, quota tracking, allocation group headers), CIL can reduce log write volume by one to two orders of magnitude.

Layer 3: Shadow Buffer Copy Amplification

CIL introduces one additional in-memory copy per commit. Each item is formatted from its live in-memory representation into a shadow buffer (xfs_log_vec.lv_buf) before being added to the CIL. This decouples the item from the log write so the item can be unlocked immediately, but it means every committed item is represented in memory at least twice: once as the live object (inode, buffer) and once as the formatted shadow. The shadow is later copied into the iclog when the CIL pushes.

Memory path for a single logged inode:

xfs_inode (in memory)
  → iop_format() → lv_buf (shadow buffer, CIL holds it)
       → xlog_write() → iclog data buffer (ring)
            → disk

Three copies before the data reaches the log device. This amplification is intentional: it removes the need to hold any lock on the live object during log I/O, enabling the parallelism that makes delayed logging viable.

Layer 4: Metadata Cascade Amplification (B-tree Fan-out)

A single application-visible operation triggers a cascade of internal metadata changes, each of which must be logged independently. The worst case occurs during extent allocation on a filesystem with all optional B-trees enabled:

Operation stepItems logged
Inode size/extent count updateXFS_LI_INODE
BMBT (extent map B-tree) blockXFS_LI_BUF × (tree height)
AGF header updateXFS_LI_BUF
Free space B-tree by block (BNOBT)XFS_LI_BUF × (split depth)
Free space B-tree by size (CNTBT)XFS_LI_BUF × (split depth)
Rmap B-tree (RMAPBT, if enabled)XFS_LI_BUF × (split depth)
Refcount B-tree (REFCBT, if enabled)XFS_LI_BUF × (split depth)
AGI header (if inode allocation)XFS_LI_BUF
Inode B-tree (INOBT/FINOBT)XFS_LI_BUF × (split depth)

A single fallocate call on a filesystem with rmapbt and refcountbt enabled can log 20–40 buffer items. Each B-tree split creates a new block that must also be logged. The reservation system accounts for this worst case, which is why per- transaction reservations are large relative to the actual bytes changed.

Layer 5: Intent/Done Record Overhead

Multi-step operations (extent free, rmap update, refcount update, attribute write) write a pair of log records — an Intent before the operation and a Done after. This ensures recovery can detect and complete partial operations. The overhead is two additional log records per complex sub-operation:

EFI (Extent Free Intent)  ← written before freeing extent
  → btree updates (AGF, BNOBT, CNTBT, RMAPBT, REFCBT)
EFD (Extent Free Done)    ← written after btree updates complete

On workloads that perform many small file deletions (e.g. log rotation, build artifact cleanup), Intent/Done pairs can account for a significant fraction of log traffic. A delete of a 100-extent file generates 100 EFI/EFD pairs plus all associated B-tree buffer logs.

Layer 6: Iclog Block Padding

Log records are written in units of 512-byte blocks and padded to the next block boundary. Small transactions that log only a few hundred bytes waste the remainder of the block. On workloads with many small transactions the padding overhead can reach 30–50% of raw log bandwidth, effectively shrinking the usable log size.

The CIL largely mitigates this by batching many small transactions into a single large checkpoint record. Padding waste is then amortized across the checkpoint rather than per-transaction.


Bottleneck Catalog

Each bottleneck is described with its root cause, how it manifests in observable metrics, and what can be done to mitigate it.

B1: Log Full — Grant Head Stall

Root cause: The write head has caught up to the tail. No physical log space remains for new transactions. All calls to xlog_grant_head_check() block on the l_write_head FIFO wait queue.

Cause chain:

Device too slow → AIL drain lags → tail does not advance
  → write head catches tail → xlog_grant_head_wait() blocks all writers

Symptoms:

  • All application threads stall in xfs_log_reserve() simultaneously — a global filesystem freeze from the application’s perspective.
  • dmesg may show XFS: xlog_grant_log_space: sleep if debug logging enabled.
  • iostat shows log device at 100% utilisation with very low metadata device I/O (metadata writes are blocked waiting for log space).
  • /proc/fs/xfs/stat: xs_trans_ail stalled; xs_push_ail_success near zero.

Mitigations:

  • Increase log size (mkfs.xfs -l size=... or external log device).
  • Move the log to a dedicated faster device (NVMe vs. HDD).
  • Reduce B-tree fan-out amplification by enabling bigtime, nrext64, or choosing a larger block size to pack more records per B-tree node.
  • Reduce the number of enabled optional B-trees if rmap/reflink are not required.

B2: CIL Context Switch Contention

Root cause: The CIL push worker acquires xc_ctx_lock as a writer to swap the live context. During this window, all concurrent xlog_cil_commit() calls block waiting for the read lock. On many-core systems with high transaction rates the context switch becomes a serialisation barrier.

Symptoms:

  • CPU profiles show many threads spinning or sleeping in xlog_cil_commit().
  • Short bursts of very high lock wait time correlate with checkpoint boundaries.
  • Transaction commit latency has a periodic spike pattern matching the CIL push interval (every few hundred milliseconds under load).

Mitigations:

  • The CIL is already tuned to minimise the write-lock hold time (context is swapped then lock released immediately). The main lever is reducing push frequency by ensuring the log is large enough that the CIL soft limit (XLOG_CIL_SPACE_LIMIT, ~12.5% of log) is not hit too often.
  • Workloads that issue many synchronous fsync calls force CIL pushes on every call. Batching fsync (e.g. using sync_file_range or application-level buffering) reduces push frequency.

B3: Iclog Spinlock Serialisation (l_icloglock)

Root cause: Every thread writing to an iclog must hold l_icloglock for the duration of the copy. The lock is a raw spinlock. On systems with 32+ CPUs all running concurrent CIL push workers, the spinlock degrades to a bottleneck.

Symptoms:

  • perf or ftrace shows high time in _raw_spin_lock called from xlog_write_iclog() or xlog_state_get_iclog_space().
  • Log write bandwidth plateaus below the device’s sequential write capacity.
  • Adding more CPUs does not improve log throughput.

Mitigations:

  • Use larger iclogs (mkfs.xfs -l version=2,size=...,su=262144 sets 256 KB iclogs). Larger iclogs mean fewer lock acquisitions per unit of log data.
  • Reduce the number of concurrent CIL pushes by ensuring workload transactions are large enough to batch well before hitting the CIL limit.

B4: Reservation Overestimation on Small Logs

Root cause: Each transaction type holds a worst-case reservation for the entire duration of the transaction, even if the actual log usage is a fraction of that. On a small log (< 256 MB), the sum of all in-flight reservations can exhaust the reserve head even when the physical log has space, causing false stalls.

Symptoms:

  • xlog_grant_head_check() blocks on l_reserve_head even though l_write_head has available space.
  • Log device utilisation is low but transaction latency is high.
  • Reducing concurrent writer count relieves the stall.

Mitigations:

  • Increase log size. Reservations are a fixed fraction of log size; a larger log accommodates more concurrent in-flight transactions.
  • Avoid small logs on high-concurrency filesystems. The minimum practical log size for a busy filesystem is typically 512 MB; 1–2 GB is common on production systems.

B5: AIL Tail Stall — Pinned Items

Root cause: Items in the AIL that are still pinned by in-flight CIL transactions cannot be flushed. If the CIL checkpoint does not complete quickly enough (e.g., because log I/O is slow), the tail cannot advance. xfsaild counts pinned items against the stuck threshold and backs off after 100 consecutive pinned items.

Symptoms:

  • /proc/fs/xfs/stat: xs_push_ail_pinned dominates over xs_push_ail_success.
  • xfsaild sleep time is consistently 20–50 ms (backoff mode).
  • ail_log_flush counter increments rapidly, causing redundant CIL flushes.
  • Log I/O latency is high (slow log device or iclog congestion).

Mitigations:

  • Faster log device reduces the time between CIL commit and iclog I/O completion, unpinning items sooner.
  • If the workload uses explicit fsync, confirm that it is not being called at a rate that prevents the CIL from batching effectively.

B6: AIL Tail Stall — Buffer Lock Contention

Root cause: xfsaild uses trylock on all buffers. Under heavy concurrent writeback from the page cache or other kernel paths, many metadata buffers are already locked when xfsaild tries to acquire them. Each failure increments the stuck counter toward the 100-item backoff threshold.

Symptoms:

  • /proc/fs/xfs/stat: xs_push_ail_locked grows alongside xs_push_ail_pinned.
  • High iowait on the metadata device during writeback storms.
  • xfsaild alternates between 0 ms (making progress) and 20 ms (backoff) with no clear pattern.

Mitigations:

  • Reduce concurrent writeback pressure via vm.dirty_background_ratio and vm.dirty_ratio.
  • On NVMe, increase the nr_requests queue depth to absorb more concurrent I/Os without serialising at the block layer.

B7: AIL Single-Threaded Traversal

Root cause: xfsaild is one thread per filesystem. The AIL traversal loop is sequential. On workloads that accumulate millions of dirty metadata items (e.g. large rsync, git clone of a large repository, database checkpoint), the traversal itself consumes significant CPU time before items are submitted.

Symptoms:

  • xfsaild CPU usage is consistently high (near 100% of one core).
  • Log device I/O queue is not saturated — xfsaild is the bottleneck, not the device.
  • Cursor restarts are frequent: xfsaild repeatedly restarts from the AIL minimum due to concurrent deletions during inode cluster flushes.

Mitigations:

  • There is no kernel-level tuning knob to parallelise xfsaild. The mitigation is to reduce the number of items in the AIL at any one time by ensuring the CIL pushes frequently and items are promptly written to disk.
  • Increasing vm.dirty_expire_centisecs delays the page cache writeback that competes with xfsaild, reducing cursor invalidation interference.

B8: Checkpoint Ordering Stall

Root cause: xlog_cil_order_write() enforces that commit records appear in strictly ascending checkpoint sequence order. A slow checkpoint (due to large checkpoint size or log I/O contention) blocks all higher-sequence checkpoints from writing their commit records, even if their data has already been written to iclogs.

Symptoms:

  • Multiple CIL push workers are stalled in xlog_cil_order_write() waiting on xc_commit_wait.
  • Log device appears idle despite pending checkpoint data.
  • Checkpoint commit latency increases proportionally to checkpoint I/O latency.

Mitigations:

  • Reduce checkpoint size by reducing the CIL soft limit (not directly tunable at runtime; requires log size adjustment since the limit is a fraction of log size).
  • Faster log device reduces per-checkpoint I/O time, shortening the ordering wait.

B9: Write Amplification from Optional B-Trees

Root cause: Enabling rmapbt (reverse mapping) and reflink (reference counting) adds two additional B-trees that must be updated on every extent allocation, deallocation, and CoW operation. Each B-tree update is a separate logged buffer item. On workloads with high extent churn, these trees double or triple the number of buffer items logged per operation.

Symptoms:

  • Log write bandwidth is significantly higher after enabling reflink or rmapbt compared to a plain filesystem.
  • Per-transaction reservation sizes are larger (visible via xfs_logprint).
  • Extent allocation operations are slower under concurrency due to higher per- transaction lock hold times.

Mitigations:

  • Do not enable rmapbt or reflink if the workload does not require them. These features cannot be disabled after mkfs.
  • Use a larger block size to increase B-tree node fanout, reducing tree height and therefore split frequency.

B10: Recovery Time from Large Logs

Root cause: Recovery replays every log record between tail_lsn and head_lsn at crash time. A larger log retains more history. A filesystem that was writing heavily immediately before the crash will have a full or nearly-full log to replay.

Symptoms:

  • Mount time is minutes rather than seconds after an unclean shutdown.
  • Recovery I/O is visible on the log device during mount.
  • dmesg shows XFS: starting recovery followed by a long gap before XFS: Ending recovery.

Mitigations:

  • Use barrier=1 (default) to ensure log records are committed before the device acknowledges the write, keeping the recovery window bounded.
  • A dedicated log device with lower write latency reduces the time to write checkpoints, keeping tail_lsn closer to head_lsn at any given moment (less to replay).
  • Do not artificially inflate log size beyond what is needed. A log larger than necessary does not improve steady-state performance and increases worst-case recovery time.

Amplification Summary

LayerWhat is amplifiedCIL mitigation
Fundamental WALEvery metadata write appears twice (log + disk)None — inherent to WAL
ReloggingHot items logged once per transactionYes — once per checkpoint
Shadow buffer copies3 in-memory copies before diskUnavoidable cost of lock-free commit
B-tree cascade10–40 buffer items per file operationPartial — items are batched per checkpoint
Intent/Done pairs2 log records per multi-step operationPartial — both records batched in checkpoint
Iclog paddingUp to 512 bytes wasted per transactionYes — padding amortised across checkpoint
Optional B-trees2× log traffic with rmapbt + reflinkNone — structural overhead

Summary of Critical Paths

PathKey bottleneck
Transaction allocationl_reserve_head FIFO wait when log is full
CIL commitxc_ctx_lock contention during context switch
CIL push → iclog writel_icloglock on every 32 KB iclog block
Iclog I/O completionBlock device latency
AIL push — device throughputxfsaild delwri queue depth vs. device bandwidth
AIL push — pinned itemsLong-running transactions pin tail; ail_log_flush triggers CIL force
AIL push — buffer contentiontrylock failures accumulate; 100-item stuck threshold triggers backoff
AIL push — single-threadedSequential traversal; cursor restarts on concurrent deletion
AIL cluster flushail_lock drop/reacquire per inode cluster; cursor restart on remove
Log tail advance__xfs_ail_assign_tail_lsn() wakes grant head waiters; stalls if AIL never drains
RecoveryLog size × write rate at crash time

Understanding these seven paths and their limiting factors is the foundation for diagnosing and resolving XFS performance problems on write-intensive workloads.


Source References

FilePurpose
fs/xfs/xfs_log.cMain log manager, iclog state machine
fs/xfs/xfs_log_cil.cDelayed logging, CIL push worker
fs/xfs/xfs_trans_ail.cAIL daemon (xfsaild), push loop, tail assignment, cursor management
fs/xfs/xfs_trans_priv.hxfs_ail and xfs_ail_cursor structure definitions
fs/xfs/xfs_inode_item.cxfs_inode_item_push(), cluster flush dispatch
fs/xfs/xfs_buf_item.cxfs_buf_item_push(), buffer trylock and delwri queue
fs/xfs/xfs_log_recover.cCrash recovery, two-pass replay
fs/xfs/xfs_log_priv.hInternal structures (xlog, xlog_in_core, xfs_cil)
fs/xfs/xfs_log.hPublic logging API
fs/xfs/libxfs/xfs_log_format.hOn-disk format (xlog_rec_header, item types)
Documentation/filesystems/xfs/xfs-delayed-logging-design.rstAuthoritative design document

XFS Dirty Log: head/tail error analysis

Context

When _check_xfs_filesystem runs after a test that triggers ENOSPC during mmap copy-on-write (e.g. generic/173), it can report two errors:

_check_xfs_filesystem: filesystem on /dev/loop1 has dirty log
_check_xfs_filesystem: filesystem on /dev/loop1 is inconsistent (r)

What the errors mean

Dirty log

_check_xfs_filesystem unmounts the filesystem and then runs xfs_logprint -t. A clean unmount should flush all dirty buffers, drain the AIL, and write a clean log marker. If the log state is still <DIRTY> after unmount, it means XFS was force-shut down before the unmount could checkpoint the log — preventing it from writing the clean marker.

Inconsistent (r)

xfs_repair -n (read-only mode) reads the raw on-disk data structures. Because the AIL never flushed the logged changes to those structures, the on-disk state is the pre-transaction state — inconsistent with what the log says should be there. xfs_repair -n cannot replay the log to reconcile them, so it flags the filesystem as inconsistent.

A plain xfs_repair (without -n) would replay the log first and would very likely find no actual corruption.

Example log output

xfs_logprint:
    data device: 0x701
    log device: 0x701 daddr: 327728 length: 131072

    log tail: 40 head: 46 state: <DIRTY>
  • Internal log — log and data are on the same device (0x701).
  • tail: 40, head: 46 — 6 blocks of live log records sit between tail and head. These records were written to the on-disk log but the corresponding buffer writes (AGF, btree blocks, inodes) never reached the data structures. On the next mount XFS would replay them.

Transactions in the log

Transaction 1 — LSN (cycle 1, block 40), tid 0x31cd442c, 11 items

ItemDetail
AGF buffer at blkno 0x1AG 0 free space manager — space was allocated or freed
3 data buffers at 0x8, 0x10, 0x28Free space or inode btree blocks modified by the allocation
Inode 0x84 (132), flags 0x5, dsize 48Core + data fork extents — a file whose extent map was being updated

Consistent with a CoW block allocation in progress: AGF modified, btree blocks updated, and the target inode’s extent map changed.

Transaction 2 — LSN (cycle 1, block 44), tid 0xa2d07cca, 2 items

ItemDetail
Inode 0x83 (131), flags 0x1, dsize 0Core only, no extent data — an inode whose data fork was cleared or reset

Inode log item anatomy

Using inode 0x84 from transaction 1 as an example:

INO: cnt:3 total:3 a:0xaaaaef94fe40 len:56 a:0xaaaaef94fef0 len:176 a:0xaaaaef94ffb0 len:48
    INODE: #regs:3   ino:0x84  flags:0x5   dsize:48
    CORE inode:
        DATA FORK EXTENTS inode data:

The inode log item is composed of 3 memory regions copied into the log:

AddressLengthContent
0xaaaaef94fe4056 bytesxfs_inode_log_format — header describing which parts of the inode were dirtied
0xaaaaef94fef0176 bytesInode core (xfs_dinode) — timestamps, size, link count, etc.
0xaaaaef94ffb048 bytesData fork extent records

flags:0x5 = XFS_ILOG_CORE (0x1) | XFS_ILOG_DEXT (0x4) — both the inode core and the data fork extent list were dirtied by this transaction.

dsize:48 — 48 bytes of extent data = 3 extents (each xfs_bmbt_rec_t is 16 bytes). These are the packed extent records describing the new block mappings for the CoW operation.

Summary of what happened

  1. The test fills the filesystem and then attempts an mmap CoW write with no free space.
  2. XFS starts a transaction: allocates CoW blocks (modifying the AGF and btree blocks) and updates the target inode’s extent map.
  3. ENOSPC is hit mid-flight. XFS force-shuts down the filesystem to avoid leaving metadata in a half-updated state.
  4. The transaction records are committed to the on-disk log (that is why xfs_logprint can see them), but the AIL never flushes the modified buffers to their actual on-disk locations — the log tail is never pushed.
  5. Force-shutdown prevents any further log writes, so the clean log marker cannot be written on unmount.
  6. xfs_logprint sees <DIRTY> and xfs_repair -n sees stale on-disk structures that do not match the logged intent.

NFS

Filesystem

NFS with KRB5

Setup KDC:

  1. Install the required packages for the KDC.
[root@kdc-server ~]# dnf install krb5-libs krb5-server krb5-workstation 
  1. Edit the /etc/krb5.conf.
# To opt out of the system crypto-policies configuration of krb5, remove the
# symlink at /etc/krb5.conf.d/crypto-policies which will not be recreated.
includedir /etc/krb5.conf.d/

[logging]
    default = FILE:/var/log/krb5libs.log
    kdc = FILE:/var/log/krb5kdc.log
    admin_server = FILE:/var/log/kadmind.log

[libdefaults]
    dns_lookup_realm = false
    ticket_lifetime = 24h
    renew_lifetime = 7d
    forwardable = true
    rdns = false
    pkinit_anchors = FILE:/etc/pki/tls/certs/ca-bundle.crt
    spake_preauth_groups = edwards25519
    dns_canonicalize_hostname = fallback
    qualify_shortname = ""
    default_realm = EXAMPLE.COM
    default_ccache_name = KEYRING:persistent:%{uid}

[realms]
 EXAMPLE.COM = {
     kdc = kdc.example.com
     admin_server = kdc.example.com
 }

[domain_realm]
 .example.com = EXAMPLE.COM
 example.com = EXAMPLE.COM
  1. Create the database using the kdb5_util.
[root@kdc-server ~]# kdb5_util create -s
  1. Set ACL in the /var/kerberos/krb5kdc/kadm5.acl. Bellow settings allows anyone with secodary admin principal to have full administrative access for example: user/admin@EXAMPLE.COM`
*/admin@EXAMPLE.COM  *
  1. Create the first principal using kadmin.local at the KDC terminal:
[root@kdc-server ~]# kadmin.local -q "addprinc user/admin"
  1. Satrt krb5kdc and kadmin.
[root@kdc-server ~]# systemctl enable --now kadmin.service krb5kdc

NFS server configuration.

  1. Install packages for kerberos client and NFS server
[root@nfs-server ~]# dnf install krb5-workstation nfs-utils 
  1. NFSv4 idmapping becomes much more important to have with Kerberos. Both the server and the clients should have the same idmapping domain configured. In the /etc/idmapd.conf set the domain to your kerberos realm.
[General]
Domain = example.com
  1. Each NFS server needs a Kerberos principal for nfs/server.fqdn to be created on the KDC, and its keys added to the server’s /etc/krb5.keytab.
[root@nfs-server ~]# kadmin -p username/admin
Password for username/admin@EXAMPLE.COM: ***********
kadmin:  addprinc -nokey nfs/nfs-server.example.com
kadmin:  addprinc -nokey host/nfs-server.example.com

kadmin:  ktadd nfs/nfs-server.example.com
kadmin:  ktadd host/nfs-server.example.com

[root@nfs-server ~]# klist -ke

[root@nfs-server ~]# klist -ke
Keytab name: FILE:/etc/krb5.keytab
KVNO Principal
---- -------------------------------------------------------------------
   1 host/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha384-192
   1 host/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha256-128) 
   1 host/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha1-96) 
   1 host/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha1-96) 
   1 host/nfs-server.example.com@EXAMPLE.COM (camellia256-cts-cmac) 
   1 host/nfs-server.example.com@EXAMPLE.COM (camellia128-cts-cmac) 
   1 host/nfs-server.example.com@EXAMPLE.COM (DEPRECATED:arcfour-hmac) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha384-192) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha256-128) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (aes256-cts-hmac-sha1-96) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (aes128-cts-hmac-sha1-96) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (camellia256-cts-cmac) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (camellia128-cts-cmac) 
   1 nfs/nfs-server.example.com@EXAMPLE.COM (DEPRECATED:arcfour-hmac) 
  1. Enable and start the gssproxy.service
[root@nfs-server ~]# systemctl enable --now gssproxy.service

NFS client configuration

  1. Install the nfs-utils and krb5-workstation as on the NFS server and create same configuration file krb5.conf.

  2. Add nfs-client to the kerberos.

[root@nfs-server ~]# kadmin -p username/admin
Password for username/admin@EXAMPLE.COM: ***********
kadmin:  addprinc -nokey host/nfs-client.example.com

kadmin:  ktadd host/nfs-client.example.com

[root@nfs-server ~]# klist -ke

Benchmarks

Benchmark of NFSv4.2 with different security context.

Environment

NFS Server and KDC:

  • OS: Fedora 40 Qemu KVM virtual machine.
    • 20 CPUs Intel® Xeon® Gold 5215 CPU @ 2.50GHz.
    • 250GB memory 100GB used for /dev/pmem0 emulation.
    • CPUs pinned on NUMA node 0 (NFS clients are pinned to NUMA node 1).
    • All configs are left in default values including nfs.conf.
    • NFS export is placed on XFS backing up /dev/pmem0 and not using DAX option.
    • Virtio isolated network.

NFS clients

  • OS: version from RHEL 7 to RHEL 9 Qemu KVM
    • 8 CPUs Intel® Xeon® Gold 5215 CPU @ 2.50GHz with 8GB of RAM.
    • Virtio isolated network.
    • Default NFS mount options:
fedora-kdc.example.com:/mnt/nfs  /mnt/sys    nfs4  rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1    0  0
fedora-kdc.example.com:/mnt/nfs  /mnt/krb5   nfs4  rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1   0  0
fedora-kdc.example.com:/mnt/nfs  /mnt/krb5i  nfs4  rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5i,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1  0  0
fedora-kdc.example.com:/mnt/nfs  /mnt/krb5p  nfs4  rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=krb5p,clientaddr=192.168.0.2,local_lock=none,addr=192.168.0.1  0  0

Base line testing with follwing fio configuration:

# cat nfs_08k_rand_write.ini
[global]
direct=1 
ioengine=libaio
bs=8k
gtod_reduce=1
size=2G 
iodepth=128
rw=randwrite
group_reporting=1
numjobs=4
filename_format=/mnt/$jobname/nfs.$jobnum

[sys]
stonewall=1
[krb5]
stonewall=1
[krb5i]
stonewall=1
[krb5p]
stonewall=1

Results on the NFS server writing directly to the local filesystem. The stats on diskstats are zero as /dev/pmem0 does not report anything to /proc/diskstats.

[root@fedora-kdc ~]# fio xfs_08k_rand_write.ini
nfs: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
fio-3.36
Starting 4 processes
Jobs: 4 (f=4): [w(4)][94.1%][w=7443MiB/s][w=953k IOPS][eta 00m:02s]
nfs: (groupid=0, jobs=4): err= 0: pid=8415: Thu Aug 15 11:22:14 2024
  write: IOPS=331k, BW=2586MiB/s (2712MB/s)(80.0GiB/31677msec); 0 zone resets
   bw (  MiB/s): min=  179, max= 7466, per=99.74%, avg=2579.39, stdev=829.78, samples=252
   iops        : min=22962, max=955700, avg=330161.65, stdev=106212.09, samples=252
  cpu          : usr=6.09%, sys=48.51%, ctx=695334, majf=0, minf=31
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=0,10485760,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
  WRITE: bw=2586MiB/s (2712MB/s), 2586MiB/s-2586MiB/s (2712MB/s-2712MB/s), io=80.0GiB (85.9GB), run=31677-31677msec

Disk stats (read/write):
  pmem0: ios=0/0, sectors=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

 

RHEL 7:

  • Encryption type: aes256-cts-hmac-sha1-96
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 0 (all jobs):
  WRITE: bw=179MiB/s (187MB/s), 179MiB/s-179MiB/s (187MB/s-187MB/s), io=8192MiB (8590MB), run=45878-45878msec

Run status group 1 (all jobs):
  WRITE: bw=173MiB/s (182MB/s), 173MiB/s-173MiB/s (182MB/s-182MB/s), io=8192MiB (8590MB), run=47316-47316msec

Run status group 2 (all jobs):
  WRITE: bw=130MiB/s (137MB/s), 130MiB/s-130MiB/s (137MB/s-137MB/s), io=8192MiB (8590MB), run=62825-62825msec

Run status group 3 (all jobs):
  WRITE: bw=81.9MiB/s (85.8MB/s), 81.9MiB/s-81.9MiB/s (85.8MB/s-85.8MB/s), io=8192MiB (8590MB), run=100064-100064msec

 

RHEL 8:

  • Encryption type: aes256-cts-hmac-sha384-192
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 1 (all jobs):
  WRITE: bw=176MiB/s (185MB/s), 176MiB/s-176MiB/s (185MB/s-185MB/s), io=8192MiB (8590MB), run=46488-46488msec

Run status group 2 (all jobs):
  WRITE: bw=154MiB/s (161MB/s), 154MiB/s-154MiB/s (161MB/s-161MB/s), io=8192MiB (8590MB), run=53239-53239msec

Run status group 3 (all jobs):
  WRITE: bw=152MiB/s (159MB/s), 152MiB/s-152MiB/s (159MB/s-159MB/s), io=8192MiB (8590MB), run=53974-53974msec

 

RHEL 9:

  • Encryption type: aes256-cts-hmac-sha384-192
sys: (g=0): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5: (g=1): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5i: (g=2): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
krb5p: (g=3): rw=randwrite, bs=(R) 8192B-8192B, (W) 8192B-8192B, (T) 8192B-8192B, ioengine=libaio, iodepth=128
...
Run status group 0 (all jobs):
  WRITE: bw=198MiB/s (208MB/s), 198MiB/s-198MiB/s (208MB/s-208MB/s), io=8192MiB (8590MB), run=41394-41394msec

Run status group 1 (all jobs):
  WRITE: bw=204MiB/s (214MB/s), 204MiB/s-204MiB/s (214MB/s-214MB/s), io=8192MiB (8590MB), run=40156-40156msec

Run status group 2 (all jobs):
  WRITE: bw=179MiB/s (187MB/s), 179MiB/s-179MiB/s (187MB/s-187MB/s), io=8192MiB (8590MB), run=45868-45868msec

Run status group 3 (all jobs):
  WRITE: bw=151MiB/s (158MB/s), 151MiB/s-151MiB/s (158MB/s-158MB/s), io=8192MiB (8590MB), run=54363-54363msec

 

Overview

OSFilesystemSECRANDOM WRITE 8k
Fedora 40XFS–2586MiB/s
RHEL 7NFSsys179MiB/s
RHEL 7NFSkrb5173MiB/s
RHEL 7NFSkrb5i130MiB/s
RHEL 7NFSkrb5p81.9MiB/s
RHEL 8NFSsys205MiB/s
RHEL 8NFSkrb5176MiB/s
RHEL 8NFSkrb5i154MiB/s
RHEL 8NFSkrb5p152MiB/s
RHEL 9NFSsys189MiB/s
RHEL 9NFSkrb5182MiB/s
RHEL 9NFSkrb5i180MiB/s
RHEL 9NFSkrb5p162MiB/s

Storage

Targetcli

Scripts for creating loopback disks

targetcli /loopback create wwn=naa.5000000000000000 \
for i in {00..16}; \
    do lvcreate -y -n disk$i -L100G export; \
    targetcli /backstores/block create dev=/dev/export/disk$i name=disk$i; \
    targetcli /backstores/block/disk$i set attribute optimal_sectors=4096; \
    targetcli /loopback/naa.5000000000000000/luns create \
    storage_object=/backstores/block/disk$i; \
done

Input Oputput I/

Breef overview of current (2023) I/O mechanism supported Linux kernel.

Synchronous

Represented by system calls read()/write(), pread()/pwrite() and it’s vectored version of readv()/writev() and preadv()/pwritev().

Asynchronous

Represented either by POSIX aio (man aio), libaio and latest io_uring.

Buffered

All IO which ends up in page cache and are not directly written to but are written by the page background process.

Direct I/O:

With direct I/O data is read from or written to the storage device (e.g., HDD, SSD, NVMe) without being cached in the operating system’s buffer cache. This means that data is transferred directly between the application’s memory and the storage device, bypassing any intermediate caching layers in the operating system.

Benefits of Direct I/O:

  1. Direct I/O is often used in scenarios where data consistency and control over I/O operations are critical, such as in databases, file systems, and some scientific computing applications.

  2. It can help ensure that data is written to and read from the storage device without interference from the operating system’s cache, which can be especially important for applications that require strict data durability or real-time performance. By bypassing the cache, direct I/O can reduce the variability in I/O response times that can occur with cached I/O.

x86_64 REGISTERS

General-Purpose Registers

The 64-bit versions of the ‘original’ x86 registers are named:
  • rax - register a extended
  • rbx - register b extended
  • rcx - register c extended
  • rdx - register d extended
  • rbp - register base pointer (start of stack)
  • rsp - register stack pointer (current location in stack, growing downwards)
  • rsi - register source index (source for data copies)
  • rdi - register destination index (destination for data copies)
The registers added for 64-bit mode are named:
  • r8 - register 8
  • r9 - register 9
  • r10 - register 10
  • r11 - register 11
  • r12 - register 12
  • r13 - register 13
  • r14 - register 14
  • r15 - register 15
These may be accessed as:
  • 64-bit registers using the r prefix: rax, r15.
  • 32-bit registers using the e prefix or d suffix: eax, r15d.
  • 16-bit registers using no prefix or a w suffix : ax, r15w.
  • 8-bit registers using h (“high byte” of 16 bits) suffix: ah, bh.
  • 8-bit registers using l (“low byte” of 16 bits) suffix or ‘b’ suffix: al, r15b.
  ================ rax (64 bits)
          ======== eax (32 bits)
              ====  ax (16 bits)
              ==    ah (8 bits)
                ==  al (8 bits)
Usage during syscall/function call:

First six arguments are in rdi, rsi, rdx, rcx, r8d, r9d; remaining arguments are on the stack. For syscalls, the syscall number is in rax. For procedure calls, rax should be set to 0. The called routine is expected to preserve rsp, rbp, rbx, r12, r13, r14 and r15 but may trampleany other registers. Return value is in rax.

GIT

Github notes.

Git hub submit changes to pull request:

  • rebase
  • force push

More verbose:

  1. changes
  2. commit,
  3. git rebase -i HEAD~2’ use ‘f’ to squash the new code into the previous one + the previous commit message

xfstests random.c bug

A Deep-Dive Compiler Forensics Case Study

1. The Context: RPM Default Flags and xfstests

The investigation began with compiling xfstests-dev, a standard Linux filesystem testing suite, using Red Hat’s default RPM build optimizations (%optflags).

Modern RPM builds inject a massive payload of performance and security flags, including:

-O2 -flto=auto -ffat-lto-objects -fexceptions -g -grecord-gcc-switches -pipe -Wall -Werror=format-security -Wp,-D_FORTIFY_SOURCE=3 -fstack-protector-strong -m64 -march=x86-64-v3 -mtune=generic ...

When building an Autotools project, these flags must be passed to the ./configure script so they are properly tested and baked into the resulting Makefile, rather than passing them directly to make (which overwrites internal project include flags like -I).

# The correct way to inject RPM flags into Autotools
./configure CFLAGS="$(rpm --eval '%{optflags}')" CXXFLAGS="$(rpm --eval '%{optflags}')" LDFLAGS="$(rpm --eval '%{__global_ldflags}')"
make

2. The Mystery: Diverging Execution

Once compiled, a bizarre issue emerged in the nametest binary. Running the exact same code with the exact same pseudo-random number generator (PRNG) seed (-s 1) produced slightly different outputs depending on the compilation flags:

Unstripped (Standard Flags):

creates: 10 OK, 0 EEXIST (10 total, 0% EEXIST)

Stripped (RPM Flags with LTO & -O2):

creates: 10 OK, 1 EEXIST (11 total, 9% EEXIST)

Because the threshold for these operations relied on op = random() % 100, this divergence proved that the internal state of the PRNG was generating different mathematical sequences despite starting with the identical seed.

3. The Red Herrings

Debugging compiler differences is notoriously difficult. Several theories were tested and ultimately ruled out:

  • Implicit Function Declarations (GCC 14): Did the strictness of GCC 14 cause ./configure to fail a feature check, silently falling back to glibc’s rand() instead of POSIX random()?

    • Ruled Out: nm output proved xfstests was statically compiling its own custom random.c implementation.
  • char Signedness: Red Hat RPM flags include -funsigned-char. Did casting bytes from a signed array alter the PRNG state?

    • Ruled Out: Recompiling with -fsigned-char did not fix the issue.
  • Missing -fwrapv: Was signed integer overflow causing the optimizer to alter the math?

    • Initial Test Failed: Passing -fwrapv to CFLAGS did not fix the issue. (Note: This was a false negative due to missing LDFLAGS, which became critical later.)

4. The Smoking Gun: Assembly Analysis

To bypass the compiler’s “black box,” the raw assembly of both binaries was dumped and compared:

objdump -d --no-show-raw-insn -M intel ./nametest-stripped > stripped.asm
objdump -d --no-show-raw-insn -M intel ./nametest-unstripped > unstripped.asm

The C code in nametest.c contained two back-to-back PRNG calls inside a loop:

ip = &table[ random() % totalnames ];
op = random() % 100;

Because of Link-Time Optimization (-flto), GCC merged random.c and nametest.c, allowing it to inline the PRNG math directly into the loop.

Looking at stripped.asm, the optimizer did something highly destructive:

1940: mov    edi,DWORD PTR [r12]       # edi = original_it
# ...
1777: lea    eax,[rdi*4+0x0]           # new_it = original_it * 4  <--- THE SMOKING GUN
177e: lea    edx,[rax-0x1]             
1781: mov    DWORD PTR [r12],eax       # save new_it to memory

The compiler had hardcoded original_it * 4 to handle both PRNG calls simultaneously, completely skipping a critical safety branch inside the random.c logic.

5. The Root Cause: Weaponized Undefined Behavior

The legacy 1994 PRNG code in random.c contained the following logic:

if (it <= 0)    
    it = (it + it) ^ MASK;
else
    it = it + it;

The original author explicitly relied on signed integer overflow. When it + it exceeded the 32-bit integer limit, it would wrap around and become a negative number. The next time the function was called, the if (it <= 0) branch would catch the negative number and apply the MASK.

How Modern GCC Handled It:

In the C standard, signed integer overflow is Undefined Behavior (UB).

When LTO inlined both random() calls, GCC’s value-range analysis looked at the second call and assumed: “If it was positive in the first call, it + it cannot mathematically be negative because signed overflow is illegal. Therefore, the < 0 branch is impossible code.”

GCC completely deleted the safety branch and simply multiplied the variable by 4. Once the seed reached 1,073,741,824, it * 4 overflowed into 0. The MASK was missed, permanently altering the mathematical sequence of the PRNG for the rest of the test.

6. The Fix

The correct way to preserve 1994 logic in a 2024 compiler pipeline is to remove the Undefined Behavior entirely. In C, unsigned integer overflow is 100% legal and defined to wrap modulo $2^n$.

By casting the variables to unsigned integers for the arithmetic, the compiler is forced to respect the wrap-around, preserving the original signed check without triggering the optimizer’s UB deletion.

The Patched Code (random.c):

it = is[0];
leh = is[1];

/* Use unsigned arithmetic to safely wrap the overflow without triggering UB */
uint32_t u_it = (uint32_t)it;

if (it <= 0)	
    u_it = (u_it + u_it) ^ MASK;
else
    u_it = u_it + u_it;
    
it = (int32_t)u_it;

With this patch, the PRNG safely survives Link-Time Optimization (-flto) and aggressive -O2 heuristics, ensuring the stripped and unstripped binaries generate the exact same random sequence.


7. Standalone Reproduction

The bug was reproduced in isolation to confirm the root cause without involving xfstests. The reproduction lives in:

randtest/
├── random.c       — PRNG implementation (_irandm, _random, random, srandom, get_it)
├── random.h       — declarations
├── randtest.c     — main: two sequential random() calls per iteration with diagnostics
└── bug_demo.c     — self-contained single-file reproduction (see below)

7.1 Two-file reproduction (random.c + randtest.c)

randtest.c calls random() twice per iteration and reads saved_seed[0] via get_it() before and after each call, so the corrupted PRNG state is visible directly:

# correct behaviour
gcc -O0 -o randtest_plain randtest.c random.c

# triggers the bug via LTO cross-file inlining
gcc $(rpm --eval '%{optflags}') $(rpm --eval '%{build_ldflags}') \
    -o randtest_rpm randtest.c random.c

Output diverges at iteration 16 — the first iteration after it crosses 2^30 and the doubled value wraps past INT32_MAX:

# plain (correct)
iter 16:  it=1187941550  r1=138120522  it=-1919084196  r2=1734299348  it=945648879   <-- OVERFLOW

# RPM (buggy)
iter 16:  it=1187941550  r1=138120522  it=-1919084196  r2=2084125751  it=456798904

r1 is still identical (both calls share the same entry state for that iteration), but r2 diverges because the inlined second call skips the MASK branch and writes a corrupted value to saved_seed[0].

7.2 Single-file reproduction (bug_demo.c)

The bug does not require LTO. A single translation unit compiled with plain -O2 is enough — the compiler inlines _irandm freely within the file and applies the same cross-call value-range analysis.

gcc -O2 -fno-lto -o bug_demo_O2 bug_demo.c   # triggers the bug
gcc -O0          -o bug_demo_O0 bug_demo.c   # correct reference

Same divergence, same iteration, no linker flags involved.

7.3 Why the printf inside _irandm masks the bug

During development a diagnostic printf was placed inside _irandm itself. This silently cured the bug: printf is an external call with side effects, which GCC treats as a full memory barrier. The optimizer can no longer track it across the call boundary, so it conservatively keeps both branches. The numbers stayed identical across builds.

The only visible artifact was that the RPM build silently dropped the "<-- OVERFLOW" label from its output — the ternary (it_old > 0 && it < 0) was statically eliminated because GCC knew (from UB reasoning in the else branch) that it < 0 is impossible after it = it + it with a positive it. Correct numbers, missing label: a subtler manifestation of the same UB exploitation.

Moving the printf to main (via get_it()) removed the barrier and restored the divergence.


8. Assembly Deep-Dive (bug_demo_O2.s)

The generated assembly (bug_demo_O2.s) shows the inlined loop. The critical section (annotated):

.L7:                                   ; loop top — load saved_seed
    leal  (%rdx,%rdx), %r8d            ; r8d = it + it  (call 1 new_it; may wrap!)
    testl %edx, %edx                   ; test ORIGINAL it (before doubling)
    jg    .L2                          ; it > 0 → fast path, NO branch check for call 2

    xorl  $593970775, %r8d             ; MASK applied (call 1, it ≤ 0 path)
    ...
    testl %r8d, %r8d                   ; call 2 gets its own branch — but only from
    jg    .L4                          ;   the it ≤ 0 entry path
    xorl  $593970775, %esi             ; MASK for call 2 if needed

.L2:                                   ; entered when original it was positive
    leal  -1(%r8), %esi                ; nit1 = (it+it) - 1
    andl  $127, %esi
    imull mt(,%rsi,4), %eax            ; leh *= mt[nit1 & 127]
    leal  0(,%rdx,4), %esi             ; THE BUG: esi = it * 4
                                        ;   GCC assumed it+it > 0 (signed overflow = UB)
                                        ;   so it*2 again needs no sign check.
                                        ;   When it = 2^30: it*4 = 2^32 → truncates to 0.
    ...
    ; falls straight into call 2 with esi = it*4, no branch, no MASK

The testl %edx, %edx / jg .L2 pair is the only branch serving both calls. It tests the pre-doubling it, not the result of the doubling (%r8d). Once jg .L2 is taken, call 2’s if (it <= 0) check is gone entirely — replaced by the hardcoded leal 0(,%rdx,4).

In contrast, -O0 emits two independent call _random instructions. Each is a black box; no cross-call value-range analysis is possible and _irandm executes its branch correctly on every invocation.

Trigger conditions summary

ScenarioBug triggers
Separate files, -O0No — no inlining
Separate files, -O2, no LTONo — compiler cannot see across files
Separate files, -O2, -fltoYes — LTO merges IR, inlines across files
Single file (bug_demo.c), -O2, no LTOYes — inlined within one translation unit
Single file, -O0No — no inlining, no value-range optimization
Any flags, printf inside _irandmNo — printf acts as a memory barrier

9. Compiler Options Reference

Flags that trigger the bug

FlagRole
-O2Enables inlining and value-range propagation — the minimum level needed
-flto=autoLink-Time Optimization: merges all translation units into one IR before optimization, giving the same cross-file view as a single .c file
-ffat-lto-objectsEmbeds both LTO IR and regular object code in each .o, so the archive works with and without LTO-aware linkers

Flags in %{optflags} relevant to this bug

-O2 -flto=auto -ffat-lto-objects -fexceptions -g -grecord-gcc-switches -pipe
-Wall -Wno-complain-wrong-lang -Werror=format-security
-Wp,-U_FORTIFY_SOURCE,-D_FORTIFY_SOURCE=3 -Wp,-D_GLIBCXX_ASSERTIONS
-specs=/usr/lib/rpm/redhat/redhat-hardened-cc1
-fstack-protector-strong
-specs=/usr/lib/rpm/redhat/redhat-annobin-cc1
-m64 -march=x86-64 -mtune=generic
-fasynchronous-unwind-tables -fstack-clash-protection
-fcf-protection -mtls-dialect=gnu2
-fno-omit-frame-pointer -mno-omit-leaf-frame-pointer

Of these, -O2 and -flto=auto are the two flags directly responsible for the bug. The rest add security hardening and debug info but do not influence the PRNG optimization.

Flags in %{build_ldflags} relevant to this bug

-Wl,-z,relro -Wl,--as-needed -Wl,-z,pack-relative-relocs -Wl,-z,now
-specs=/usr/lib/rpm/redhat/redhat-hardened-ld
-specs=/usr/lib/rpm/redhat/redhat-hardened-ld-errors
-specs=/usr/lib/rpm/redhat/redhat-annobin-cc1
-Wl,--build-id=sha1

The hardened-ld specs activate the LTO linker plugin. Without passing LDFLAGS correctly, -flto in CFLAGS compiles LTO IR into the objects but the link step discards it — which is why early -fwrapv tests appeared to fix the issue (the LTO IR was never linked).

Flags that suppress the bug

FlagEffect
-O0Disables inlining entirely; each _irandm call is a real function call
-fno-ltoDisables LTO; cross-file inlining impossible
-fwrapvTells GCC that signed integer overflow wraps (two’s complement); UB assumption removed, branch preserved
-fno-strict-overflowWeaker form of -fwrapv; disables overflow-based optimizations
__attribute__((noinline)) on _irandmPrevents inlining of the function; forces a real call boundary

Security hardening visible in the RPM binary (unrelated to the bug)

FeatureFlagEffect on binary
PIE-specs=redhat-hardened-cc1ELF type changes from EXEC to DYN
Full RELRO-Wl,-z,relro + BIND_NOWAll GOT entries made read-only after startup
BIND_NOW-Wl,-z,nowAll symbols resolved at load time
FORTIFY_SOURCE=3-D_FORTIFY_SOURCE=3printf replaced by __printf_chk, buffer overflows detected at runtime
Stack clash protection-fstack-clash-protectionProbe stack pages on allocation to prevent stack-clash attacks
CF protection (CET)-fcf-protectionendbr64 inserted at every indirect-jump target
Frame pointers kept-fno-omit-frame-pointerEnables reliable stack unwinding in profilers and crash dumps
Debug info-g -grecord-gcc-switchesDWARF sections embedded; binary grows from 13K to 19K

2’s Complement

The dominant way modern hardware represents signed integers.

Core Idea

For an N-bit integer, a negative number -x is stored as 2^N - x.

For 8-bit (N=8):
  -1   →  2^8 - 1   =  255  =  0xFF  =  11111111
  -2   →  2^8 - 2   =  254  =  0xFE  =  11111110
  -128 →  2^8 - 128 =  128  =  0x80  =  10000000

Bit Layout (8-bit)

Bit pattern  | Unsigned | Signed (2's complement)
-------------|----------|------------------------
0000 0000    |    0     |    0
0000 0001    |    1     |    1
0111 1111    |   127    |   127       ← INT_MAX
1000 0000    |   128    |  -128       ← INT_MIN (sign bit flips)
1000 0001    |   129    |  -127
1111 1110    |   254    |   -2
1111 1111    |   255    |   -1

The sign bit (MSB) being 1 means negative. The range is asymmetric: one more negative value than positive.

How to Negate Manually

Two equivalent methods:

Method 1: Flip all bits, then add 1

 5  =  0000 0101
~5  =  1111 1010   (flip)
-5  =  1111 1011   (add 1)

Method 2: 2^N - x

-5  =  256 - 5  =  251  =  1111 0101   ✓ (same result)

Why Hardware Loves It

Addition and subtraction use the same circuit for signed and unsigned:

  0000 0101  (+5)
+ 1111 1011  (-5 in 2's complement)
-----------
1 0000 0000  → carry discarded → 0000 0000 = 0  ✓

No special subtraction hardware needed. This is the primary reason 2’s complement won over alternatives like sign-magnitude or 1’s complement.

The Three Historical Alternatives (Mostly Dead)

SchemeHow -5 looks (8-bit)Problem
Sign-magnitude1000 0101Two zeros (+0 and -0), complex arithmetic
1’s complement1111 1010Also two zeros, end-around carry needed
2’s complement1111 1011One zero, simple arithmetic ✓

Reading a Bit Pattern

Example: 1111 1111

Method 1 — sign bit formula (MSB has weight -2^(N-1), rest are normal):

1111 1111
│└──────┘
│  positional values: 64+32+16+8+4+2+1 = 127
│
└─ sign bit: -128

Total: -128 + 127 = -1

Method 2 — flip and add 1:

1111 1111  → flip → 0000 0000  → add 1 → 0000 0001 = 1

Magnitude is 1, sign bit is 1, so the value is -1.

Verify:

  1111 1111  (-1)
+ 0000 0001  (+1)
-----------
1 0000 0000  → carry dropped → 0  ✓

All-ones is always -1 in 2’s complement, regardless of bit width (8, 16, 32, 64).


Example: 1000 0000

Method 1 — sign bit formula:

1000 0000
│└──────┘
│  positional values: 0+0+0+0+0+0+0 = 0
│
└─ sign bit: -128

Total: -128 + 0 = -128

Method 2 — flip and add 1:

1000 0000  → flip → 0111 1111  → add 1 → 1000 0000

You get 1000 0000 back — it’s its own negation. This is why -128 has no positive counterpart in 8-bit signed: +128 doesn’t fit (0111 1111 = 127 is the max).

The asymmetry:

INT_MIN = -128  =  1000 0000
INT_MAX = +127  =  0111 1111

|INT_MIN| > INT_MAX — this is why abs(INT_MIN) is undefined behavior in C. Negating -128 would require +128, which overflows.

Same Bits, Different Meaning

1000 0000 represents different values depending on interpretation:

InterpretationValue
Unsigned128
2’s complement signed-128

The hardware stores 1000 0000 — whether that’s 128 or -128 is decided purely by how your code declares the variable:

uint8_t u = 0x80;   // 128
int8_t  s = 0x80;   // -128

Casting uint32_t to int32_t

When you cast uint32_t v to int32_t:

  • The bit pattern does not change
  • The CPU just reinterprets bit 31 as a sign bit
  • If bit 31 is 1 (value ≥ 2^31), the result is negative
uint32_t v = 0x80000000;  // 2147483648, bit 31 set
int32_t  s = (int32_t)v;  // -2147483648 — same bits, signed interpretation

In C11/C17 this is implementation-defined behavior. In C23 it is finally guaranteed by the standard, as 2’s complement is now mandated for all signed integer types.

Overflow Wraps the Number Line into a Circle

Visualize it as a clock:

         0
    -1       1
  -2           2
         ...
   -128      127   (8-bit)
        -127

Adding past INT_MAX wraps to INT_MIN, and vice versa:

result = (a + b) mod 2^N   (then reinterpret as signed)

Kernel Build

Applying old config to new kernel

Copy existing .config to new kernel source tree:

cp kvm-vm-ext4-xfs.config /path/to/linux-7.x/.config
cd /path/to/linux-7.x

All new options set to NO (tightest result)

make KCONFIG_ALLCONFIG=.config allnoconfig

Keeps every option from old config, sets all new/unknown options to n. Kconfig auto-resolves dependencies — if existing CONFIG_VIRTIO_PCI=y now depends on new CONFIG_FOO, it gets force-enabled.

All new options set to their defaults

make olddefconfig

Usually n, but some may default to y. Less strict than allnoconfig.

List what is new

make listnewconfig

Prints every option old config doesn’t cover. Useful to scan before blanket-disabling.

Full recipe

cd /path/to/linux-7.x
cp /path/to/kvm-vm-ext4-xfs.config .config
make listnewconfig > new_options.txt
make KCONFIG_ALLCONFIG=.config allnoconfig
make -j$(nproc)

Building kernel src.rpm with mock

Install

sudo dnf install mock rpm-build
sudo usermod -aG mock $USER
newgrp mock

Build from source RPM

mock -r fedora-43-x86_64 --rebuild kernel-6.19.0-1.src.rpm

Results in /var/lib/mock/fedora-43-x86_64/result/.

Build from spec + sources

rpmbuild -bs kernel.spec \
  --define "_sourcedir $(pwd)" \
  --define "_srcrpmdir $(pwd)"

mock -r fedora-43-x86_64 --rebuild kernel-*.src.rpm

The spec expects linux.tar.gz as Source0 and config as Source1. Both must exist in same directory.

Useful mock flags

# Keep build tree for debugging
mock --no-clean --rebuild kernel-*.src.rpm

# Use specific config (EPEL, CentOS Stream, etc.)
mock -r centos-stream+epel-9-x86_64 --rebuild kernel-*.src.rpm

# Inject .config before build
mock --init
mock --copyin kvm-vm-ext4-xfs.config /builddir/build/SOURCES/config
mock --no-clean --rebuild kernel-*.src.rpm

# Speed up with tmpfs (needs RAM)
mock --enable-plugin=tmpfs --rebuild kernel-*.src.rpm

List available mock configs

ls /etc/mock/*.cfg

Controlling CPU count in mock

Default: all CPUs. Verify with:

rpm --eval '%_smp_mflags'

Override via mock define

mock -r fedora-43-x86_64 \
  --define "_smp_mflags -j16" \
  --rebuild kernel-*.src.rpm

Override via rpmbuild opts

mock -r fedora-43-x86_64 \
  --rpmbuild-opts="--define '_smp_mflags -j16'" \
  --rebuild kernel-*.src.rpm

Persistent mock config

In /etc/mock/custom.cfg:

config_opts['macros']['%_smp_mflags'] = '-j16'

Limit via cgroup

systemd-run --scope -p AllowedCPUs=0-15 \
  mock -r fedora-43-x86_64 --rebuild kernel-*.src.rpm

Quick reference

CPUs desiredFlag
All-j$(nproc)
16-j16
Half-j$(( $(nproc) / 2 ))

Docs

Nextcloud internals

Cleaning of the bruteforce IP

use nextcloud;
show tables;
select * from oc_bruteforce_attempts;
delete from oc_bruteforce_attempts where IP="xxx.xxx.xxx.xxx";

Fedora KickStarts

fedora-37 /dev/vda

Ansible VM Provisioning Guide

Automated Fedora VM provisioning on KVM hypervisors using Ansible.

Inventory Configuration

[kvmhosts]
hw-server-1.intra.herbolt.com

Playbooks Overview

System Management

PlaybookPurpose
get_facts.ymlGather system info (hostname, kernel, IP, RAM, updates, EOL)
update_packages.ymlUpdate packages and restart affected services

VM Provisioning

PlaybookPurpose
provision_fedora_vm_home.ymlFedora VMs on home server (UEFI, static IP, bridge)
remove_vm_home.ymlRemove VMs from home server with safety checks

get_facts.yml

Gathers system information with switchable output format.

# Human-readable (default)
ansible-playbook -i inventory.ini get_facts.yml

# JSON output
ansible-playbook -i inventory.ini get_facts.yml -e "output_format=json"

update_packages.yml

Update packages and restart affected services.

Tags:

  • upgrade-all — all packages
  • upgrade-security — security only
  • upgrade-kernel — kernel only
  • services — restart affected services
  • check-upgrades — check without installing
# Update all + restart services
ansible-playbook -i inventory.ini update_packages.yml --tags upgrade-all,services

# Security updates only
ansible-playbook -i inventory.ini update_packages.yml --tags upgrade-security

# Check what needs updating
ansible-playbook -i inventory.ini update_packages.yml --tags check-upgrades

# Update all but don't restart services
ansible-playbook -i inventory.ini update_packages.yml \
  --tags upgrade-all \
  -e "auto_restart_services=false"

provision_fedora_vm_home.yml

What It Does

  1. Detects Fedora version (latest-1 stable)
  2. Creates VM index (auto-increments if name exists)
  3. Creates LVM storage
  4. Generates kickstart configuration
  5. Provisions VM with virt-install
  6. Configures bridge network with static IP
  7. Sets up SSH key authentication

Default Specifications

VM Name:        fedora-vm-1 (auto-incremented)
Fedora:         Latest-1 (e.g., Fedora 41 if 42 is latest)
Target:         hw-server-1.intra.herbolt.com
Boot:           UEFI without Secure Boot
Memory:         16GB RAM
CPUs:           12 vCPUs
Disk:           20GB LV (virtio, io_uring, writeback cache)
Filesystem:     XFS on LVM
Swap:           zram (8GB compressed in-memory)
Network:        Bridge (br0) with static IP (172.168.31.250/24)
Gateway:        172.168.31.1
DNS:            172.168.31.30
Security:       SELinux enforcing, SSH keys, firewall enabled

Quick Start

# Default VM
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml

# Custom VM
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=myvm" \
  -e "vm_memory=8192" \
  -e "vm_vcpus=4"

Configuration Variables

VM Configuration

VariableDefaultDescription
vm_base_namefedora-vmBase name (index auto-appended)
vm_memory16096RAM in MB
vm_vcpus12Virtual CPU cores
vm_disk_size20Disk size in GB
vm_secure_bootfalseEnable UEFI Secure Boot

Storage Configuration

VariableDefaultDescription
vg_namevmVolume group name
vm_filesystemxfsFilesystem: xfs, ext4, btrfs

Network Configuration

VariableDefaultDescription
vm_network_typebridgebridge or macvtap
vm_bridge_namebr0Bridge name
vm_use_static_iptrueStatic IP enabled
vm_static_ip172.168.31.250Static IP address
vm_static_netmask255.255.255.0Network mask
vm_static_gateway172.168.31.1Default gateway
vm_static_dns172.168.31.30DNS server

User Configuration

VariableDefaultDescription
vm_root_passwordfedora123Root password
vm_root_ssh_keyssh-ed25519 AAAA...SSH public key
vm_userfedoraRegular user name
vm_user_passwordfedora123Regular user password
vm_timezoneEurope/PragueTimezone
vm_keyboardusKeyboard layout
fedora_versionauto-detectedFedora version (latest-1)

Command Examples

Resource Allocation

# Minimal VM
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=minimal" \
  -e "vm_memory=2048" \
  -e "vm_vcpus=2" \
  -e "vm_disk_size=10"

# Standard VM
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=standard" \
  -e "vm_memory=8192" \
  -e "vm_vcpus=4" \
  -e "vm_disk_size=50"

# High-performance VM
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=performance" \
  -e "vm_memory=32768" \
  -e "vm_vcpus=16" \
  -e "vm_disk_size=200"

# Database server
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=database" \
  -e "vm_memory=65536" \
  -e "vm_vcpus=24" \
  -e "vm_disk_size=500" \
  -e "vm_filesystem=xfs"

Network Configuration

# Override static IP
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_static_ip=172.168.31.251"

# Multiple VMs with sequential IPs
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=web" \
  -e "vm_static_ip=172.168.31.101"

ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=web" \
  -e "vm_static_ip=172.168.31.102"

# Use different bridge
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_bridge_name=virbr0"

# Use macvtap instead of bridge
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_network_type=macvtap" \
  -e "vm_macvtap_interface=eno1"

Custom static IP configuration

ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_static_ip=10.0.0.100" \
  -e "vm_static_netmask=255.255.255.0" \
  -e "vm_static_gateway=10.0.0.1" \
  -e "vm_static_dns=8.8.8.8"

Security

# Enable Secure Boot
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_secure_boot=true"

# Random passwords
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_root_password=$(openssl rand -base64 32)" \
  -e "vm_user_password=$(openssl rand -base64 32)"

# Custom SSH key
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_root_ssh_key=$(cat ~/.ssh/id_ed25519.pub)"

# Maximum security
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_base_name=secure" \
  -e "vm_secure_boot=true" \
  -e "vm_root_password=$(openssl rand -base64 32)" \
  -e "vm_user_password=$(openssl rand -base64 32)"

Force specific Fedora version

ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "fedora_version=40"

Filesystem Options

XFS (Default)

/boot/efi  - EFI partition (600MB)
/boot      - XFS (1GB)
/          - XFS on LVM (10GB)
/home      - XFS on LVM (4GB)
swap       - zram (8GB compressed)

Best for large files, databases, high-performance workloads.

ext4

ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_filesystem=ext4"

Same layout with ext4. Best for maximum compatibility, general purpose.

Btrfs

ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  -e "vm_filesystem=btrfs"

Layout uses subvolumes (@root, @home). Best for snapshots, compression, copy-on-write.


Tags

TagPurpose
alwaysCore initialization
preparationInstall packages, setup
packagesEnsure dependencies
servicesEnable libvirtd
storageCreate LVM storage
checkValidation checks
isoDownload Fedora ISO
kickstartCreate kickstart file
networkNetwork configuration
vmVM provisioning
infoDisplay information
# Pre-flight check only
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  --tags check,info

# Prepare everything but don't provision
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  --tags preparation,storage,iso,kickstart

# Just provision (assumes prep done)
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  --tags vm

Batch VM Creation

# Create 5 web servers
for i in {1..5}; do
  ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
    -e "vm_base_name=web-server"
done

# Web server farm with sequential IPs
for i in {1..3}; do
  IP="172.168.31.$((100+i))"
  ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
    -e "vm_base_name=web" \
    -e "vm_memory=16384" \
    -e "vm_vcpus=8" \
    -e "vm_disk_size=100" \
    -e "vm_static_ip=${IP}"
done

Reusable Configuration Files

dev-vm.yml:

vm_base_name: dev
vm_memory: 4096
vm_vcpus: 2
vm_disk_size: 20
vm_filesystem: btrfs
vm_secure_boot: false

prod-vm.yml:

vm_base_name: prod
vm_memory: 32768
vm_vcpus: 16
vm_disk_size: 500
vm_filesystem: xfs
vm_secure_boot: true
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml -e "@dev-vm.yml"
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml -e "@prod-vm.yml"

remove_vm_home.yml

Remove VMs including storage with safety checks (10-second countdown).

# List all VMs
ansible-playbook -i inventory.ini remove_vm_home.yml

# Remove specific VMs
ansible-playbook -i inventory.ini remove_vm_home.yml \
  -e 'vms_to_remove=["fedora-vm-1","fedora-vm-2"]'

# Filter by base name
ansible-playbook -i inventory.ini remove_vm_home.yml \
  -e 'vm_filter=fedora-vm'

What gets removed: VM definition from libvirt, disk (logical volume), force-stops running VMs.


VM Management

After Provisioning

# Connect to console
virsh console fedora-vm-1
# Exit with: Ctrl + ]

# Get VM IP
virsh domifaddr fedora-vm-1

# SSH (key-based auth configured)
ssh root@172.168.31.250

# Add to /etc/hosts
echo "172.168.31.250  fedora-vm-1" >> /etc/hosts

Lifecycle Commands

virsh list --all              # List all VMs
virsh start fedora-vm-1       # Start
virsh shutdown fedora-vm-1    # Graceful stop
virsh destroy fedora-vm-1     # Force stop
virsh reboot fedora-vm-1      # Restart
virsh autostart fedora-vm-1   # Autostart on boot
virsh dominfo fedora-vm-1     # VM info

# Delete VM and disk
virsh destroy fedora-vm-1
virsh undefine fedora-vm-1
lvremove /dev/vm/fedora-vm-1-boot

# LVM snapshot
lvcreate -L 5G -s -n fedora-vm-1-boot-snap /dev/vm/fedora-vm-1-boot

Troubleshooting

“Unknown OS name ‘fedora42’” — Automatic: playbook uses closest available version.

“Bridge ‘br0’ does not exist” — Create bridge first or use different one: -e "vm_bridge_name=virbr0"

“Volume Group ‘vm’ does not exist” — Create VG: vgcreate vm /dev/sdb or use different: -e "vg_name=storage"

VM already exists — Auto-increments index. To replace: destroy old VM first.

Can’t SSH to VM:

virsh list                      # Is it running?
virsh domifaddr fedora-vm-1     # Get IP
ping 172.168.31.250             # Can you reach it?
ssh -v root@172.168.31.250      # Debug SSH

Debug Commands

# Dry run
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml --check

# Verbose
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml -vvv

# Step through
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml --step

# Show config without running
ansible-playbook -i inventory.ini provision_fedora_vm_home.yml \
  --tags info \
  -e "vm_base_name=test"

Quick Reference

SettingDefault
RAM16GB
CPUs12
Disk20GB
FilesystemXFS
NetworkBridge (br0), Static IP (172.168.31.250/24)
Userfedora / fedora123
RootSSH key + fedora123
Swapzram (8GB compressed)
BootUEFI without Secure Boot
SELinuxEnforcing
Servicessshd, chronyd, qemu-guest-agent
InstallationNetwork install from download.fedoraproject.org

Requirements

Control Node

  • Ansible 2.9+
  • Python 3.6+

Hypervisor (hw-server-1)

  • Fedora with libvirt and KVM
  • LVM volume group named ‘vm’
  • Bridge (br0) configured
  • Python 3
  • Internet connectivity for network installation

Packages (Auto-installed)

  • libvirt, libvirt-daemon-kvm
  • virt-install
  • qemu-guest-agent
  • edk2-ovmf (UEFI firmware)
  • lvm2

Infrastructure

Home lab running on a single Dell PowerEdge R620 hypervisor (hw-server-1) with KVM virtualization. All services run as VMs on a private bridge network (172.16.31.0/24), with the hypervisor handling NAT and port forwarding to the public IP.

Network Topology

Internet
  │
  ▼
hw-server-1 (5.59.97.199) ── eno1, public zone, masquerade
  │
  ├── br0 (172.16.31.1/24) ── trusted zone
  │     │
  │     ├── http-server-1  (.10) ── Nginx, Nextcloud, MariaDB
  │     ├── mail-server-1  (.20) ── Postfix, Dovecot, LDAP
  │     ├── vpn-server-1   (.30) ── WireGuard, BIND DNS
  │     ├── clp-1          (.40) ── Node.js (Cleaning Plan app)
  │     └── pod-server-1   (.50) ── Docker (Nextcloud AppAPI HARP)
  │
  └── virbr0 (192.168.122.1/24) ── libvirt default (down)

Port Forwarding (public IP)

PortsDestinationService
80, 443/tcp172.16.31.10Web (http-server-1)
25, 143, 465, 587, 993/tcp172.16.31.20Mail (mail-server-1)
5252, 51820, 51821/udp172.16.31.30WireGuard (vpn-server-1)

Server Summary

ServerOSvCPUsRAMRole
hw-server-1Fedora 4324 (physical)188 GiBKVM hypervisor
http-server-1Fedora 42 (EOL)1232 GiBWeb server
mail-server-1Fedora 42 (EOL)24 GiBMail server
vpn-server-1Fedora 42 (EOL)44 GiBVPN + DNS
clp-1Fedora 42 (EOL)24 GiBNode.js app (Cleaning Plan)
pod-server-1Fedora 42 (EOL)84 GiBContainer host (Nextcloud AppAPI)

Common Issues Across Fleet

  • Fedora 42 EOL — http-server-1, mail-server-1, vpn-server-1, clp-1, pod-server-1 all past end-of-life
  • Trusted firewall zones — most VMs have their interface in trusted zone (ACCEPT all); pod-server-1 has firewalld masked entirely
  • Deprecated QEMU machine type — 3 VMs still on pc-q35-7.0 (http-server-1, mail-server-1, vpn-server-1)

clp-1

Role: Node.js web application server (Cleaning Plan) | Location: clp-1.intra.herbolt.com Sosreport: 2026-07-30

System

KeyValue
OSFedora Linux 42 (Server Edition)
Kernel6.17.8-200.fc42.x86_64
PlatformKVM virtual machine (QEMU Q35, edk2 UEFI)
CPU2x Intel Xeon E5-2630L v2 @ 2.40GHz
RAM3.8 GiB
Swap1.9 GiB (zram)
TimezoneEurope/Prague (CEST)

Storage

DeviceSizeMountFilesystem
vda1600M/boot/efivfat
vda21G/bootxfs
vg_system-lv_root10G/xfs
vg_system-lv_home4G/homexfs
zram01.9G[SWAP]swap

VG vg_system has 4.41 GiB free (unallocated PE).

Network

InterfaceIPZone
enp1s0172.16.31.40/24FedoraServer (default)
lo127.0.0.1/8–
  • Gateway: 172.16.31.1
  • DNS: 172.16.31.30 (via NetworkManager), local stub via systemd-resolved
  • NTP: 2.fedora.pool.ntp.org (chronyd, synchronized)

Cleaning Plan Application

The server runs a custom Node.js web application as a systemd service.

KeyValue
Servicecleaning_plan.service
Node.jsv22.20.0
Userdherbolt
ExecStart/usr/bin/node /home/dherbolt/www/cleaning-plan/server/index.js
Port3001/tcp
NODE_ENVproduction
Restarton-failure

The service is enabled via multi-user.target.wants and was running at sosreport time.

Key Services

ServiceStatusPurpose
cleaning_plan.servicerunningNode.js web app (port 3001)
sshd.servicerunningSSH access
cockpit.socketlisteningWeb management console (port 9090)
firewalld.servicerunningFirewall
chronyd.servicerunningNTP time sync
qemu-guest-agent.servicerunningKVM guest agent
systemd-resolved.servicerunningDNS resolver
systemd-oomd.servicerunningOOM killer
auditd.servicerunningSecurity audit logging
rsyslog.servicerunningSystem logging

Issues

  1. Fedora 42 reached EOL on 2026-05-27 – no security updates for 2+ months

    Upgrade to Fedora 43 immediately.

    sudo dnf upgrade --refresh
    sudo dnf install dnf-plugin-system-upgrade
    sudo dnf system-upgrade download --releasever=43
    sudo dnf system-upgrade reboot
    
  2. ModemManager.service running – unnecessary on a headless server VM

    sudo systemctl disable --now ModemManager.service
    sudo systemctl mask ModemManager.service
    
  3. pcscd.service running – Smart Card Daemon unnecessary on a headless server VM

    sudo systemctl disable --now pcscd.service pcscd.socket
    sudo systemctl mask pcscd.service pcscd.socket
    
  4. abrt-xorg.service running – Xorg log watcher unnecessary on a headless server

    sudo systemctl disable --now abrt-xorg.service
    sudo systemctl mask abrt-xorg.service
    
  5. No tuned profile active – VM is not using the virtual-guest performance profile

    sudo dnf install tuned
    sudo systemctl enable --now tuned
    sudo tuned-adm profile virtual-guest
    
  6. Hostname set to clp-1.localdomain instead of proper FQDN

    sudo hostnamectl set-hostname clp-1.intra.herbolt.com
    
  7. cleaning_plan.service missing WorkingDirectory directive

    The service unit does not set WorkingDirectory, which means the Node.js app runs with / as its working directory. If the application uses relative paths, this can cause subtle failures.

    sudo systemctl edit cleaning_plan.service
    

    Add:

    [Service]
    WorkingDirectory=/home/dherbolt/www/cleaning-plan/server
    

    Then:

    sudo systemctl daemon-reload
    sudo systemctl restart cleaning_plan.service
    

Security Notes

  • SSH PasswordAuthentication enabled (default) – disable and use key-only auth:

    echo 'PasswordAuthentication no' | sudo tee /etc/ssh/sshd_config.d/10-no-password.conf
    sudo systemctl reload sshd
    
  • SSH PermitRootLogin is prohibit-password (default) – consider disabling entirely:

    echo 'PermitRootLogin no' | sudo tee -a /etc/ssh/sshd_config.d/10-no-password.conf
    sudo systemctl reload sshd
    
  • SELinux is enforcing (targeted) – good, no action needed

  • Crypto policy is DEFAULT – consider hardening:

    sudo update-crypto-policies --set DEFAULT:NO-SHA1
    
  • Cockpit web console accessible on port 9090 – verify that access is restricted to trusted networks; cockpit is using socket activation so it only starts on demand

  • Firewall zone FedoraServer allows: ssh, cockpit, dhcpv6-client, port 3001/tcp – appropriate for this server’s role

Performance Notes

  • VG vg_system has 4.41 GiB unallocated – available for extending / or /home as needed
  • zram swap (1.9 GiB compressed) is appropriate for a KVM VM with 3.8 GiB RAM
  • No tuned profile is active; installing tuned with virtual-guest profile would optimize I/O and CPU scheduling for a VM workload
  • Node.js process consuming ~106 MB RSS – reasonable for a web application
  • irqbalance is running – appropriate for a 2-vCPU VM

Sosreport Plugins

  • sos version: 4.10.0
  • RPMs installed: 706

Loaded plugins (87): abrt, alternatives, anaconda, anacron, ata, auditd, block, boot, btrfs, cgroups, chrony, cifs, cockpit, console, coredump, cron, crypto, date, dbus, devicemapper, devices, dnf, dracut, filesys, firewall_tables, firewalld, fwupd, grub2, gssproxy, hardware, host, hts, i18n, iscsi, jars, kdump, kernel, keyutils, krb5, kvm, ldap, libraries, libvirt, login, logrotate, logs, lvm2, md, memory, multipath, networking, networkmanager, nfs, nis, nodejs, nss, ntb, openhpi, openssl, pam, pci, perl, process, processor, psacct, python, release, rpm, samba, scsi, selinux, services, smartcard, sos_extras, soundcard, ssh, sssd, sudo, sunrpc, system, systemd, sysvipc, teamd, tpm2, udev, udisks, unbound, unpackaged, usb, wireless, x11, xen, xfs

Notable: tuned plugin not loaded (tuned package is not installed on this system)

hw-server-1

Role: KVM hypervisor | Location: hw-server-1.intra.herbolt.com Sosreport: 2026-07-30

System

FieldValue
OSFedora Linux 43
Kernel7.1.3-101.fc43.x86_64
HardwareDell PowerEdge R620
CPUIntel Xeon E5-2630L v2 @ 2.40GHz (24 logical cores, 2 sockets)
RAM188 GiB
Swap8 GiB zram
TimezoneEurope/Prague

Storage

System disk (sda, 446.6 GiB)

PartitionSizeMountFilesystem
sda1600 MB/boot/efivfat
sda21 GB/bootxfs
sda335 GiB/ext4 (LVM os/root)
sda4400 GiB(PV in VG vm)unallocated

VM storage disk (sdb, 3.3 TiB) — VG: vm

Logical VolumeSizeVMStatus
http_server_1_boot30 GiBhttp-server-1open
http_server_1_data_0~2 TiBhttp-server-1open
mail_server_1_boot15 GiBmail-server-1open
mail_server_1_data_025 GiBmail-server-1open
vpn_server_1_boot20 GiBvpn-server-1open
clp-1-boot20 GiBclp-1open
pod-server-1-boot10 GiBpod-server-1open
fedora-csb-1-boot50 GiBfedora-csb-1closed
fedora-vm-1-boot20 GiBfedora-vm-1closed
git-1-boot20 GiBgit-1closed

VG vm free space: 1.47 TiB

Network

InterfaceIPZonePurpose
eno15.59.97.199/26publicPublic-facing NIC
br0172.16.31.1/24trustedVM bridge network
virbr0192.168.122.1/24libvirtDefault libvirt NAT (down)
vnet0-vnet4--VM tap interfaces

Default gateway: 5.59.97.193 via eno1

Firewall

Public zone (eno1)

  • Masquerade enabled (NAT for VMs)
  • Ports: 50000-60000/tcp+udp
  • SSH: restricted via rich rules (CZ ipset + 2 specific IPs)
  • Port forwarding:
PortsDestinationService
80, 443/tcp172.16.31.10http-server-1 (web)
25, 143, 465, 993, 587/tcp172.16.31.20mail-server-1 (mail)
5252, 51820, 51821/udp172.16.31.30vpn-server-1 (WireGuard)

Trusted zone (br0)

Target: ACCEPT (full trust for VM bridge traffic)

Virtual Machines

VMStatevCPUsRAMAutostart
http-server-1running1232 GiByes
mail-server-1running24 GiByes
vpn-server-1running44 GiByes
clp-1running24 GiByes
pod-server-1running84 GiByes
fedora-csb-1shut off48 GiBno
fedora-vm-1shut off1215.7 GiBno
git-1shut off215.7 GiBno

Running totals: 28 vCPUs (overcommitted vs 24 logical), 48 GiB RAM of 188 GiB

Key Services

  • virtqemud, virtlxcd (modular libvirt daemons)
  • sshd, firewalld, NetworkManager, chronyd
  • auditd, crond, smartd, lm_sensors
  • pmcd, pmie, pmlogger (PCP monitoring)
  • SELinux: Enforcing
  • Failed units: None

Issues

  1. Deprecated QEMU machine type — http-server-1, mail-server-1, vpn-server-1 still use pc-q35-7.0 (clp-1 upgraded to pc-q35-9.1; pod-server-1, fedora-csb-1, fedora-vm-1, git-1 upgraded to pc-q35-10.1)

    # For each affected VM, update machine type (requires VM shutdown)
    virsh shutdown <vm-name>
    virsh edit <vm-name>
    # Change: <type arch='x86_64' machine='pc-q35-7.0'>
    # To:     <type arch='x86_64' machine='pc-q35-10.1'>  (matches latest on this host)
    # Verify available types:
    virsh domcapabilities | grep -oP 'machine=.*?pc-q35[^"]*' | sort -V | tail -1
    virsh start <vm-name>
    
  2. vCPU overcommit — 28 vCPUs vs 24 logical cores (acceptable under current load, monitor only)

Security Notes

  • SSH geo-restricted — good: only CZ ipset + 2 specific IPs can reach port 22

  • SELinux enforcing — good baseline

  • X11 forwarding enabled — unnecessary on production hypervisor (set via /etc/ssh/sshd_config.d/50-redhat.conf)

    sed -i 's/^X11Forwarding yes/X11Forwarding no/' /etc/ssh/sshd_config.d/50-redhat.conf
    systemctl reload sshd
    
  • Port range 50000-60000 open — wide TCP+UDP range on public zone. Audit and restrict:

    # Check what actually listens in that range
    ss -tlnp | awk '$4 ~ /:5[0-9]{4}/'
    # Remove if unused
    firewall-cmd --permanent --zone=public --remove-port=50000-60000/tcp
    firewall-cmd --permanent --zone=public --remove-port=50000-60000/udp
    firewall-cmd --reload
    
  • No intrusion detection — no fail2ban for SSH brute-force

    dnf install -y fail2ban
    cat > /etc/fail2ban/jail.local << 'EOF'
    [sshd]
    enabled = true
    maxretry = 5
    bantime = 3600
    EOF
    systemctl enable --now fail2ban
    
  • No automatic security updates

    dnf install -y dnf-automatic
    # Edit /etc/dnf/automatic.conf: apply_updates = yes, upgrade_type = security
    sed -i 's/^apply_updates.*/apply_updates = yes/' /etc/dnf/automatic.conf
    sed -i 's/^upgrade_type.*/upgrade_type = security/' /etc/dnf/automatic.conf
    systemctl enable --now dnf-automatic.timer
    

Performance Notes

  • No tuned profile detected — set tuned-adm profile virtual-host for KVM hypervisor workloads (optimizes CPU governor, I/O scheduler, transparent hugepages)
  • zram swap only — fine at current 25% RAM utilization. If VMs scale up, consider disk-backed swap as fallback
  • LVM writeback cache — verify VM disk LVs use cache=writeback in virt-install for I/O performance (at cost of crash safety)
  • PCP monitoring running — good for historical performance data

Sosreport Plugins

Enabled extras: tuned, ipmitool, numa, kvm, chrony

All recommended plugins for this server role are already enabled. SMART disk data is collected by the hardware plugin (enabled by default).

sos report -e tuned,ipmitool,numa,kvm,chrony

http-server-1

Role: Web server (Nginx + PHP-FPM + MariaDB) | Location: http-server-1.intra.herbolt.com Sosreport: 2026-07-30

System

FieldValue
OSFedora Linux 42 (EOL 2026-05-27)
Kernel6.19.14-108.fc42.x86_64
PlatformQEMU/KVM (Q35)
CPU12 vCPUs — Intel Xeon E5-2630L v2 @ 2.40GHz
RAM32 GiB
Swap8 GiB zram
TimezoneEurope/Prague

Storage

DeviceSizeMountFilesystem
vda1200 MB/boot/efivfat
vda21 GB/bootXFS
vg00-rootvol28.3 GB/XFS (LVM)
vg00-swapvol500 MB-swap (unused, zram used)
vdb2 TB/var/www/html/nextcloud/dataXFS

VG vg00: fully allocated (0 free PE). Unused swap LV wastes 500 MB.

Network

InterfaceIPZone
enp3s0172.16.31.10/24trusted

Gateway: 172.16.31.1 | DNS: 172.16.31.30

Static routes via 172.16.31.30: 172.16.11.0/24, 172.16.10.0/24, 10.103.1.12/32, 172.16.40.0/24

Web Stack

Nginx 1.30.1

Global: worker_processes auto, keepalive_timeout 8192s, proxy timeouts 300s

VhostBackendPurpose
data.herbolt.comPHP-FPM (Nextcloud)Nextcloud instance
lukas.herbolt.comStatic (mdbook) + /galleryPersonal site
cockpit.herbolt.comproxy 127.0.0.1:9090Cockpit web console
byty.herbolt.comproxy 172.16.31.40:3001App proxy
mail.herbolt.comRoundcubeWebmail
piwigo.herbolt.comPHP-FPMPhoto gallery
cert.local.lcStaticCertificate files
nvr.herbolt.comProxyNVR
recorder.herbolt.comProxyRecorder
video.herbolt.comProxyVideo
status.herbolt.comProxy/staticStatus page
svjvodova51-53.czProxy/staticSVJ site

All SSL via Let’s Encrypt (Certbot 3.3.0). HTTP/2 enabled on SSL vhosts (byty, cert, cockpit, lukas, nextcloud, roundcube, video).

Disabled configs: etherpad, ldap, owncloud, tracker (*.bck)

PHP-FPM 8.4.21

Serves Nextcloud, Roundcube, Piwigo.

MariaDB 10.11.16

DatabasePurpose
nextcloudPrimary app
roundcubemailWebmail
piwigoPhoto gallery
owncloudLegacy (unused)
etherpadLegacy (unused)

Valkey 8.0.9

Redis-compatible key-value store on port 6379, localhost only. Unix socket at /run/valkey/valkey.sock. Used by Nextcloud for caching/locking.

Key Services

ServicePurpose
nginxReverse proxy + web server
php-fpmPHP FastCGI
mariadbDatabase
valkeyCache (Nextcloud)
nextcloud-push-notificationPush daemon
nextcloud-cron.timerNextcloud periodic tasks
nextcloud-preview-gen.timerNextcloud preview generator
cockpitWeb management
pmcd/pmloggerPCP monitoring
SELinuxPermissive
Failed unitsNone

Issues

  1. Fedora 42 EOL — support ended 2026-05-27, no security updates

    # Check available upgrade path
    dnf install -y dnf-plugin-system-upgrade
    dnf system-upgrade download --releasever=43
    dnf system-upgrade reboot
    
  2. Legacy databases — owncloud, etherpad still in MariaDB (nginx configs disabled)

    # Backup before dropping
    mysqldump owncloud > /root/owncloud_backup.sql
    mysqldump etherpad > /root/etherpad_backup.sql
    # Drop if confirmed unused
    mysql -e "DROP DATABASE owncloud;"
    mysql -e "DROP DATABASE etherpad;"
    
  3. keepalive_timeout 8192s — ~2.3 hours, risk of connection exhaustion

    # In /etc/nginx/nginx.conf, change to:
    # keepalive_timeout 120;
    sed -i 's/keepalive_timeout 8192/keepalive_timeout 120/' /etc/nginx/nginx.conf
    nginx -t && systemctl reload nginx
    
  4. Unused swap LV — vg00-swapvol 500 MB allocated but zram used instead

    lvremove /dev/vg00/swapvol
    lvextend -l +100%FREE /dev/vg00/rootvol
    xfs_growfs /
    

Security Notes

  • SELinux permissive — critical on public-facing web server

    # Check what would break in enforcing mode
    ausearch -m AVC -ts recent | audit2why
    # Switch to enforcing (immediate)
    setenforce 1
    # Make permanent
    sed -i 's/^SELINUX=permissive/SELINUX=enforcing/' /etc/selinux/config
    
  • SSH PermitRootLogin yes

    sed -i 's/^PermitRootLogin yes/PermitRootLogin prohibit-password/' /etc/ssh/sshd_config
    systemctl reload sshd
    
  • Firewall effectively open — enp3s0 in trusted zone (ACCEPT all)

    # Move interface to public zone with explicit services
    firewall-cmd --permanent --zone=trusted --remove-interface=enp3s0
    firewall-cmd --permanent --zone=public --add-interface=enp3s0
    firewall-cmd --permanent --zone=public --add-service={http,https,ssh}
    firewall-cmd --reload
    
  • SSH password auth + GSSAPI + X11 forwarding enabled — PasswordAuthentication defaults to yes (not explicitly set), GSSAPIAuthentication and X11Forwarding set in drop-in /etc/ssh/sshd_config.d/50-redhat.conf

    # Disable password auth (add explicit setting before crypto-policies include)
    sed -i '1i PasswordAuthentication no' /etc/ssh/sshd_config
    # Disable GSSAPI and X11 in drop-in
    sed -i 's/^GSSAPIAuthentication yes/GSSAPIAuthentication no/' /etc/ssh/sshd_config.d/50-redhat.conf
    sed -i 's/^X11Forwarding yes/X11Forwarding no/' /etc/ssh/sshd_config.d/50-redhat.conf
    systemctl reload sshd
    
  • No fail2ban

    dnf install -y fail2ban
    cat > /etc/fail2ban/jail.local << 'EOF'
    [sshd]
    enabled = true
    maxretry = 5
    bantime = 3600
    
    [nginx-http-auth]
    enabled = true
    maxretry = 5
    bantime = 3600
    EOF
    systemctl enable --now fail2ban
    

Performance Notes

  • keepalive_timeout 8192s — reduce to 60-120s. Current value holds connections open ~2.3 hours, wasting worker slots under load
  • worker_connections 512 — low for 12 vCPUs serving multiple vhosts. Consider increasing to 1024-2048
  • Valkey on localhost — good: Nextcloud caching avoids MariaDB round-trips. Unix socket also available for lower latency
  • HTTP/2 enabled — configured on most SSL vhosts (byty, cert, cockpit, lukas, nextcloud, roundcube, video). Piwigo, status, and svjvodova are HTTP-only
  • MariaDB tuning — innodb_buffer_pool_size=12G (good for 32 GiB RAM), innodb_buffer_pool_instances=10, query_cache_size=256M, slow query log enabled (long_query_time=1s). innodb_flush_log_at_trx_commit=2 trades durability for performance
  • No opcache tuning visible — PHP opcache settings affect Nextcloud performance significantly. Check opcache.memory_consumption and opcache.interned_strings_buffer

Sosreport Plugins

Enabled extras: nginx, mysql, redis, tuned, selinux, valkey, cockpit

Note: No php or certbot sos plugins exist. PHP-FPM config collected via filesys/services. Cert info via openssl (enabled by default).

mail-server-1

Role: Mail server (Postfix + Dovecot + LDAP) | Location: mail-server-1.intra.herbolt.com Sosreport: 2026-07-30

System

FieldValue
OSFedora Linux 42 (EOL 2026-05-27)
Kernel6.19.14-108.fc42.x86_64
PlatformQEMU/KVM (Q35)
CPU2 vCPUs — Intel Xeon E5-2630L v2 @ 2.40GHz
RAM3.8 GiB
Swap3.8 GiB zram
TimezoneEurope/Prague
Tunedvirtual-guest

Storage

DeviceSizeMountFilesystem
sda1600 MB/boot/efivfat
sda21 GB/bootXFS
fedora_fedora-root13.4 GB/XFS (LVM)
sdb (LABEL=vmail)25 GB/homeXFS (lazytime)
zram03.8 GBswap-

LVM VG fedora_fedora: fully allocated. /home holds all virtual mailboxes.

Network

InterfaceIPZone
enp3s0172.16.31.20/24trusted

Gateway: 172.16.31.1 | DNS: 172.16.31.30, 1.1.1.1

Static routes via 172.16.31.30: 172.16.10.0/24, 172.16.40.0/24, 172.16.11.0/24

Mail Configuration

Postfix (MTA) — v3.9.1

SettingValue
myhostnamemx0.herbolt.com
mydomainherbolt.com
inet_interfacesall
inet_protocolsipv4
message_size_limit30 MB
TLSLet’s Encrypt (mail.herbolt.com), opportunistic
SASL authDovecot-based, TLS-only

Virtual mailbox domains: herbolt.com, gastro-horovice.cz, weboveaplikace.net

Mailboxes (15 total):

  • herbolt.com: lukas, danek, daniel, info, katka, vaclav, tomik, test, dana, fani, amelia
  • gastro-horovice.cz: herbolt, shop
  • weboveaplikace.net: dherbolt, lherbolt

Anti-spam pipeline:

  1. SPF checks (policyd-spf)
  2. Greylisting (Postgrey)
  3. DMARC (OpenDMARC milter)
  4. SpamAssassin (spamass-milter)

Note: OpenDKIM is installed and enabled but not configured in smtpd_milters — outgoing mail is not DKIM-signed. See Issues.

Outbound relay: not configured (relayhost is empty). Orphaned /etc/postfix/sasl_passwd file exists with credentials for 37.46.208.54 — see Security Notes.

Dovecot (IMAP/POP3/LMTP) — v2.3.21.1

SettingValue
Protocolsimap, pop3, lmtp, submission, sieve
SSLrequired
Authpasswd-file (/etc/dovecot/passwd)
Mail locationmaildir:~/Maildir (under /home/vmail/vhosts/)
Sieveenabled, global spam filter
ManageSieveport 4190
Submission relay127.0.0.1:25 (local Postfix)
Max connections/user50

389 Directory Server (LDAP)

Instance: local-lc — running. Backend uses deprecated BDB (should migrate to MDB).

Fail2Ban

JailStatus
postfixActive
dovecotActive
sieveActive

Key Services

ServiceStatus
postfixrunning
dovecotrunning
postgreyrunning
opendmarcrunning
spamassassinrunning
fail2banrunning
dirsrv@local-lcrunning
opendkimFAILED
SELinuxEnforcing

Issues

  1. Fedora 42 EOL — no security updates since 2026-05-27

    dnf install -y dnf-plugin-system-upgrade
    dnf system-upgrade download --releasever=43
    dnf system-upgrade reboot
    
  2. opendkim.service FAILED — service fails to start, and opendkim socket is not configured in Postfix smtpd_milters — outgoing mail is not DKIM-signed, hurts deliverability

    # Check why it failed
    systemctl status opendkim
    journalctl -u opendkim --no-pager -n 50
    # Common fix: key file permissions
    chown opendkim:opendkim /etc/opendkim/keys/ -R
    chmod 0600 /etc/opendkim/keys/*/default.private
    systemctl restart opendkim
    # After fixing the service, add opendkim to Postfix milter chain
    postconf -e "smtpd_milters = unix:/var/run/opendkim/opendkim.sock, unix:/var/run/opendmarc/opendmarc.sock, unix:/var/run/spamass-milter/postfix/sock"
    postconf -e "non_smtpd_milters = unix:/var/run/opendkim/opendkim.sock"
    systemctl reload postfix
    # Verify DKIM signing works
    opendkim-testkey -d herbolt.com -s default -vvv
    
  3. 389-DS TLS certificates expired — both Self-Signed-CA and Server-Cert

    # Check current cert expiry
    dsctl local-lc tls show-cert Server-Cert
    # Generate new self-signed certs
    dsconf local-lc security certificate del --name "Self-Signed-CA"
    dsconf local-lc security certificate del --name "Server-Cert"
    dsconf local-lc security ca-certificate generate --self-sign --name "Self-Signed-CA"
    dsconf local-lc security certificate generate --ca "Self-Signed-CA" --name "Server-Cert"
    systemctl restart dirsrv@local-lc
    
  4. 389-DS BDB backend deprecated — migrate to MDB (LMDB)

    dsctl local-lc stop
    dsctl local-lc db2ldif --replication userRoot
    dsctl local-lc dblib-bdb2mdb
    dsctl local-lc start
    # Verify
    dsconf local-lc backend config get | grep nsslapd-backend-implement
    
  5. certbot not installed — Let’s Encrypt cert renewal at risk

    dnf install -y certbot
    # Check existing cert expiry
    openssl x509 -enddate -noout -in /etc/letsencrypt/live/mail.herbolt.com/fullchain.pem
    # Set up auto-renewal
    systemctl enable --now certbot-renew.timer
    
  6. /home at 95% — mailbox delivery will fail when full

    # Check largest mailboxes
    du -sh /home/vmail/vhosts/herbolt.com/*/ | sort -rh | head
    # On hypervisor: extend the data disk LV
    lvextend -L +25G /dev/vm/mail_server_1_data_0
    # Back on mail-server-1: grow filesystem
    xfs_growfs /home
    
  7. Unnecessary services running

    systemctl disable --now avahi-daemon avahi-daemon.socket
    systemctl disable --now ModemManager
    

Security Notes

  • Dovecot passwords stored in PLAIN text — /etc/dovecot/passwd contains cleartext passwords

    # Generate hashed password for each user
    doveadm pw -s BLF-CRYPT -p "password_here"
    # Replace {PLAIN}password with {BLF-CRYPT}$2y$... in /etc/dovecot/passwd
    # Update dovecot auth scheme
    # In /etc/dovecot/conf.d/auth-passwdfile.conf.ext:
    # args = scheme=BLF-CRYPT /etc/dovecot/passwd
    systemctl restart dovecot
    
  • Firewall trusted zone — all traffic accepted on all ports

    # Create mail-specific zone
    firewall-cmd --permanent --new-zone=mail
    firewall-cmd --permanent --zone=mail --add-service={smtp,smtps,imap,imaps,pop3s,ssh}
    firewall-cmd --permanent --zone=mail --add-port={587/tcp,4190/tcp}
    firewall-cmd --permanent --zone=trusted --remove-interface=enp3s0
    firewall-cmd --permanent --zone=mail --add-interface=enp3s0
    firewall-cmd --reload
    
  • PermitRootLogin yes

    sed -i 's/^PermitRootLogin yes/PermitRootLogin prohibit-password/' /etc/ssh/sshd_config
    systemctl reload sshd
    
  • Orphaned SASL relay credentials — /etc/postfix/sasl_passwd contains plaintext credentials for 37.46.208.54 but no relayhost is configured in Postfix. Either remove the file or properly configure the relay.

    # Option A: remove orphaned credentials
    rm /etc/postfix/sasl_passwd /etc/postfix/sasl_passwd.db
    # Option B: if relay is needed, configure it properly
    postconf -e "relayhost = [37.46.208.54]"
    postconf -e "smtp_sasl_auth_enable = yes"
    postconf -e "smtp_sasl_password_maps = hash:/etc/postfix/sasl_passwd"
    postconf -e "smtp_sasl_security_options = noanonymous"
    chmod 0600 /etc/postfix/sasl_passwd /etc/postfix/sasl_passwd.db
    chown root:root /etc/postfix/sasl_passwd /etc/postfix/sasl_passwd.db
    systemctl reload postfix
    
  • No SMTP rate limiting

    # Add to /etc/postfix/main.cf
    postconf -e "smtpd_client_message_rate_limit = 100"
    postconf -e "smtpd_client_connection_rate_limit = 30"
    postconf -e "anvil_rate_time_unit = 60s"
    systemctl reload postfix
    

DNS Mail Authentication Audit (verified 2026-07-29 via 1.1.1.1)

herbolt.com

RecordStatusValue
MXOK10 mx0.herbolt.com
SPFOKv=spf1 ip4:37.46.208.54/24 ip4:5.59.97.199 a -all
DKIMMISSINGdefault._domainkey resolves to wildcard CNAME (poison-ivy.herbolt.com) — no DKIM TXT record
DMARCMISSING_dmarc resolves to wildcard CNAME — no DMARC policy
MTA-STSMISSINGNo _mta-sts TXT record
TLS-RPTMISSINGNo _smtp._tls TXT record
DANE/TLSAMISSINGNo TLSA record for _25._tcp.mx0.herbolt.com
PTROK5.59.97.199 → mx0.herbolt.com (forward-confirmed)
TLS certOKLet’s Encrypt, valid until 2026-10-18, CN=mail.herbolt.com
Relay PTRMISSING37.46.208.54 has no reverse PTR record

Root cause: wildcard DNS. herbolt.com has a wildcard *.herbolt.com → poison-ivy.herbolt.com CNAME. This catches _dmarc.herbolt.com, default._domainkey.herbolt.com, _mta-sts.herbolt.com etc., making them resolve to the wildcard instead of returning NXDOMAIN or proper TXT records. DKIM and DMARC records must be created as explicit entries to override the wildcard.

gastro-horovice.cz

RecordStatusValue
MXOK10 mail.herbolt.com (→ 5.59.97.199 via wildcard)
SPFWEAKv=spf1 include:herbolt.com ~all (softfail, should be -all)
DKIMMISSINGWildcard catches default._domainkey — no DKIM TXT record
DMARCMISSINGWildcard catches _dmarc — no DMARC policy

weboveaplikace.net

RecordStatusValue
MXOK10 mail.herbolt.com (→ 5.59.97.199 via wildcard)
SPFWEAKv=spf1 include:herbolt.com ~all (softfail, should be -all)
DKIMMISSINGWildcard catches default._domainkey — no DKIM TXT record
DMARCMISSINGWildcard catches _dmarc — no DMARC policy

DNS Recommendations

1. Fix DKIM records (all 3 domains) — requires fixing opendkim service first, then publishing keys:

# On mail-server-1: generate DKIM keys if missing
opendkim-genkey -D /etc/opendkim/keys/herbolt.com -d herbolt.com -s default -b 2048
opendkim-genkey -D /etc/opendkim/keys/gastro-horovice.cz -d gastro-horovice.cz -s default -b 2048
opendkim-genkey -D /etc/opendkim/keys/weboveaplikace.net -d weboveaplikace.net -s default -b 2048

# Get the DNS TXT records to publish
cat /etc/opendkim/keys/herbolt.com/default.txt
cat /etc/opendkim/keys/gastro-horovice.cz/default.txt
cat /etc/opendkim/keys/weboveaplikace.net/default.txt

# Add explicit DNS TXT records for default._domainkey.<domain>
# IMPORTANT: on herbolt.com the wildcard CNAME will catch this
#            unless an explicit record is created to override it

2. Add DMARC records (all 3 domains) — start with monitoring, move to reject:

# DNS TXT records to add (explicit, overrides wildcard):
_dmarc.herbolt.com          TXT "v=DMARC1; p=quarantine; rua=mailto:lukas@herbolt.com; pct=100"
_dmarc.gastro-horovice.cz   TXT "v=DMARC1; p=quarantine; rua=mailto:lukas@herbolt.com; pct=100"
_dmarc.weboveaplikace.net   TXT "v=DMARC1; p=quarantine; rua=mailto:lukas@herbolt.com; pct=100"

# After verifying DKIM works and reports look clean, change p=quarantine to p=reject

3. Harden SPF on secondary domains — change ~all to -all:

gastro-horovice.cz   TXT "v=spf1 include:herbolt.com -all"
weboveaplikace.net   TXT "v=spf1 include:herbolt.com -all"

4. Add PTR for relay IP — 37.46.208.54 has no reverse DNS. Contact the relay provider to set PTR to mx0.herbolt.com or the relay’s EHLO hostname. Missing PTR causes deliverability issues with strict receivers (Gmail, Microsoft).

5. Consider MTA-STS — enforces TLS for inbound mail delivery:

# DNS TXT record:
_mta-sts.herbolt.com  TXT "v=STSv1; id=20260729"

# Publish policy at https://mta-sts.herbolt.com/.well-known/mta-sts.txt:
version: STSv1
mode: testing
mx: mx0.herbolt.com
max_age: 86400

6. Consider TLS-RPT — receive reports about TLS delivery failures:

_smtp._tls.herbolt.com  TXT "v=TLSRPTv1; rua=mailto:tls-reports@herbolt.com"

Verification script

#!/bin/bash
# Mail auth check — run from any host, queries public DNS
for domain in herbolt.com gastro-horovice.cz weboveaplikace.net; do
  echo "=== $domain ==="
  echo "MX:      $(dig +short @1.1.1.1 MX "$domain")"
  echo "SPF:     $(dig +short @1.1.1.1 TXT "$domain" | grep -i spf || echo 'MISSING')"
  echo "DMARC:   $(dig +short @1.1.1.1 TXT "_dmarc.$domain" | grep -i dmarc || echo 'MISSING')"
  echo "DKIM:    $(dig +short @1.1.1.1 TXT "default._domainkey.$domain" | grep -i 'v=DKIM' || echo 'MISSING')"
  echo "MTA-STS: $(dig +short @1.1.1.1 TXT "_mta-sts.$domain" | grep -i sts || echo 'MISSING')"
  echo "TLS-RPT: $(dig +short @1.1.1.1 TXT "_smtp._tls.$domain" | grep -i tls || echo 'MISSING')"
  echo
done

Performance Notes

  • 2 vCPUs for mail+LDAP+SpamAssassin — SpamAssassin is CPU-intensive. If spam volume is high, consider increasing to 4 vCPUs or offloading to rspamd (lower resource usage than SpamAssassin)
  • 3.8 GiB RAM — adequate for current mailbox count (15), but 389-DS + SpamAssassin + Postfix + Dovecot compete for memory. Monitor OOM killer (systemd-oomd running)
  • maildir format — good for concurrent access and backup. No performance concern at current scale
  • lazytime mount on /home — good: reduces inode timestamp write I/O for maildir access patterns
  • Postgrey greylisting — adds delivery delay for first-time senders. If user complaints, consider switching to rspamd greylisting with auto-whitelisting
  • LVM fully allocated — no room for root expansion without adding PV. /home (vmail) on raw disk — cannot be extended without adding another disk

Sosreport Plugins

Enabled extras: postfix, dovecot, ds, fail2ban

All recommended plugins for this server role are already enabled. No sos plugins exist for SpamAssassin, certbot, or OpenDKIM — their config is collected via general filesys/services plugins.

sos report -e postfix,dovecot,ds,fail2ban

Note: Dovecot passwd file captured by sosreport contains plaintext passwords. Consider excluding /etc/dovecot/passwd from future runs with --mask or switching to hashed passwords first.

pod-server-1

Role: Container host (Docker/Podman + Nextcloud AppAPI HARP proxy) | Location: pod-server-1.intra.herbolt.com Sosreport: 2026-07-30

System

FieldValue
OSFedora Linux 42 (Server Edition) – EOL 2026-05-13
Kernel6.14.0-63.fc42.x86_64
PlatformKVM virtual machine (QEMU Q35) on hw-server-1
CPUIntel Xeon E5-2630L v2 @ 2.40GHz (8 vCPUs)
RAM3.8 GiB
Swap1.9 GiB zram
TimezoneEurope/Prague

Storage

System disk (vda, 10 GiB)

PartitionSizeMountFilesystem
vda1600 MB/boot/efivfat
vda21 GB/bootext4
vda38.4 GiB/ext4 (LVM vg_system/lv_root)

No swap partition – swap is via zram0 (1.9 GiB).

Root filesystem at 79% (6.1 GiB used of 8.2 GiB). Container images and overlays share this space under /var/lib/containerd and /var/lib/containers.

Network

InterfaceIPPurpose
enp1s0172.16.31.50/24Primary NIC
docker0172.17.0.1/16Docker bridge (no active containers attached)

Default gateway: 172.16.31.1 via enp1s0 DNS: 172.16.31.30

Listening ports

PortProcessNotes
22/tcpsshdSSH
8780/tcphaproxy (container)Nextcloud AppAPI HARP proxy, bound to 172.16.31.50
8782/tcpfrps (container)FRP server tunnel
23000/tcpfrps (container)FRP server control
24000/tcpfrps (container)FRP server
9090/tcpsystemd (cockpit)Cockpit web console
5355/tcpsystemd-resolvedLLMNR

Container Platform

Engines

EngineVersionStatus
Docker (moby-engine)29.1.3running, enabled via socket activation
containerd2.0.7running
Podman5.7.1installed, no running containers

Docker runs with --selinux-enabled. Containerd is the backend runtime for Docker (moby).

Running containers (Docker/containerd)

Two containers running via Docker, both part of the Nextcloud AppAPI HARP stack:

Container 1 – HARP proxy (image: ghcr.io/nextcloud/nextcloud-appapi-harp:release)

  • haproxy (ports 8780 ExApps proxy, 8200 internal API, 9600 internal)
  • haproxy_agent.py (Python management agent)
  • frps (FRP server on ports 8782, 23000, 24000)
  • frpc (FRP client, reverse tunnel via /frpc-docker.toml)

Container 2 – ExApp worker

  • python3 main.py (Nextcloud ExApp backend)
  • frpc (FRP client via /frpc.toml)

Podman

Podman 5.7.1 is installed with netavark/aardvark-dns networking backend. No containers, pods, or volumes are active. One image is cached:

ImageTagSize
ghcr.io/nextcloud/nextcloud-appapi-harprelease211 MB

Container networking

  • Docker bridge: 172.17.0.0/16 (docker0, no active attachment)
  • Podman bridge: 10.88.0.0/16 (podman0, unused)
  • Docker manages its own iptables/nftables NAT and forwarding rules

Key Services

ServiceStatus
docker.servicerunning, enabled
containerd.servicerunning
sshd.servicerunning, enabled
cockpit.socketenabled (port 9090)
chronyd.servicerunning, enabled
auditd.servicerunning, enabled
rsyslog.servicerunning, enabled
pmcd.servicerunning, enabled (PCP)
pmlogger.servicerunning, enabled
systemd-resolved.servicerunning, enabled
qemu-guest-agent.servicerunning, enabled
firewalld.servicemasked
ModemManager.servicerunning, enabled
bluetooth.serviceenabled

Issues

  1. Fedora 42 is end-of-life (EOL 2026-05-13). The system reports OS Support Expired: 2month 2w 3d. No security updates are available.

    # Upgrade to Fedora 43
    dnf system-upgrade download --releasever=43
    dnf system-upgrade reboot
    
  2. SELinux is permissive at runtime but configured as enforcing. Config file says enforcing, but sestatus shows Current mode: permissive. Someone ran setenforce 0 manually or the system auto-relabeled. Containers are running with SELinux labels (container_t) but policy is not being enforced.

    # Verify no denials first
    ausearch -m AVC --start today
    # Re-enable enforcing
    setenforce 1
    # Confirm it persists across reboot (config already says enforcing)
    grep ^SELINUX= /etc/selinux/config
    
  3. firewalld is masked. The host firewall is completely disabled. Docker manages its own iptables chains for container networking, but the host INPUT chain policy is ACCEPT with no rules – all ports are exposed to the network.

    # Unmask and start firewalld
    systemctl unmask firewalld
    systemctl enable --now firewalld
    # Note: Docker will automatically integrate with firewalld via the docker zone
    # Verify docker zone is configured
    firewall-cmd --get-active-zones
    
  4. Root filesystem at 79% with only 1.7 GiB free on a 10 GiB disk. Container images and layers consume space under /var/lib/containerd and /var/lib/containers. Pulling additional images or container growth could exhaust the disk.

    # Check space consumers
    du -sh /var/lib/containerd /var/lib/containers /var/lib/docker 2>/dev/null
    # Prune unused Docker resources
    docker system prune -a
    # Prune unused Podman resources
    podman system prune -a
    # Consider expanding vg_system/lv_root if space is available on hw-server-1
    
  5. ModemManager is running. Unnecessary on a server VM with no modem hardware.

    systemctl disable --now ModemManager
    
  6. Bluetooth service is enabled. Unnecessary on a KVM virtual machine.

    systemctl disable bluetooth
    
  7. Hostname is not FQDN. Set to pod-server-1.localdomain instead of pod-server-1.intra.herbolt.com.

    hostnamectl set-hostname pod-server-1.intra.herbolt.com
    

Security Notes

  • No host firewall. firewalld is masked; iptables INPUT policy is ACCEPT with zero rules. All services (SSH, Cockpit, PCP, LLMNR) are reachable from any network. Unmask firewalld immediately.

    systemctl unmask firewalld
    systemctl enable --now firewalld
    
  • SELinux permissive. Policy violations are logged but not blocked. Any container escape or compromised process runs unconstrained. Switch to enforcing after reviewing AVC denials.

    setenforce 1
    
  • SSH allows password authentication via PAM. The drop-in 50-redhat.conf sets KbdInteractiveAuthentication no but UsePAM yes still permits PAM-based password auth. No explicit PasswordAuthentication no is set.

    echo 'PasswordAuthentication no' > /etc/ssh/sshd_config.d/10-no-password.conf
    systemctl reload sshd
    
  • SSH root login not explicitly disabled. Default Fedora behavior allows root login with keys, but this should be explicit.

    echo 'PermitRootLogin prohibit-password' > /etc/ssh/sshd_config.d/10-no-root-password.conf
    systemctl reload sshd
    
  • Cockpit (port 9090) is listening on all interfaces. Accessible from any network since firewalld is masked.

  • Container image policy is insecureAcceptAnything. Any image from any registry is accepted without signature verification (/etc/containers/policy.json).

  • Crypto policy is DEFAULT. Consider tightening if only modern clients connect.

    update-crypto-policies --set DEFAULT:NO-SHA1
    

Performance Notes

  • 8 vCPUs with 3.8 GiB RAM is adequate for the current HARP proxy workload (haproxy + Python agent + FRP tunnels + one ExApp worker using ~255 MB RSS).

  • Root filesystem on a single 10 GiB virtio disk via LVM. No separate volume for container storage – all container layers share / which limits headroom.

  • No tuned profile is installed or active. For a container workload on a VM, the virtual-guest profile would be appropriate.

    dnf install tuned
    systemctl enable --now tuned
    tuned-adm profile virtual-guest
    
  • zram swap (1.9 GiB) is configured and unused, which is expected under normal load.

  • NFS client libraries are installed (rpc_pipefs mounted, nfs-client.target enabled) but no NFS mounts are active. If NFS is not needed, disable it.

    systemctl disable nfs-client.target
    

Sosreport Plugins

Version: sos 4.8.2 Command: sos report --batch --label pod-server-1 --tmp-dir /tmp/sosreport_staging

Enabled (83 plugins): abrt, alternatives, anaconda, anacron, ata, auditd, block, boot, btrfs, cgroups, chrony, cifs, cockpit, console, containerd, containers_common, coredump, cron, crypto, date, dbus, devicemapper, devices, dnf, dracut, filesys, firewall_tables, firewalld, fwupd, grub2, gssproxy, hardware, host, hts, i18n, iscsi, jars, kdump, kernel, keyutils, krb5, kvm, ldap, libraries, libvirt, login, logrotate, logs, lvm2, md, memory, multipath, networking, networkmanager, nfs, nis, nss, ntb, openhpi, openssl, pam, pci, pcp, perl, podman, process, processor, psacct, python, release, rpm, samba, scsi, selinux, services, smartcard, sos_extras, soundcard, ssh, sssd, sudo, sunrpc, system, systemd, sysvipc, teamd, udev, udisks, unbound, unpackaged, usb, wireless, x11, xen

Notable for this server: containerd and podman plugins ran and collected container data. No docker sos plugin exists (Docker data comes via containerd plugin and filesystem collection).

Missing but relevant: tuned plugin did not run because tuned is not installed.

vpn-server-1

Role: VPN gateway + DNS server | Location: vpn-server-1.intra.herbolt.com Sosreport: 2026-07-30

System

FieldValue
OSFedora Linux 42 (EOL 2026-05-27)
Kernel6.19.14-108.fc42.x86_64
PlatformQEMU/KVM (Q35)
CPU4 vCPUs — Intel Xeon E5-2630L v2 @ 2.40GHz
RAM3.5 GiB
Swap3.5 GiB zram
TimezoneEurope/Prague

Storage

DeviceSizeMountFilesystem
vda1200 MB/boot/efivfat
vda21 GB/bootXFS
vg00-rootvol5.9 GB/XFS (LVM)
zram03.5 GBswap-

LVM VG vg00: ~13 GiB free (rootvol expandable).

Network

InterfaceTypeIPZone
enp3s0Ethernet172.16.31.30/24trusted
srv0WireGuard (NM)172.16.40.1/24wireguard-srv
usr0WireGuard (NM)172.16.100.1/24wireguard-usr
wg0WireGuard (wg-quick)172.168.200.1/24trusted

Gateway: 172.16.31.1 | DNS: 127.0.0.1 (local BIND)

DNS

BIND 9.21.20 (bind9-next) running as local resolver. All VMs use 172.16.31.30 as DNS.

WireGuard Tunnels

srv0 — Server-to-server (port 5252, NM-managed)

Local: 172.16.40.1/24

PeerAllowed IPsKeepalive
Peer 1172.168.40.2/3225s
Peer 2172.16.40.3/32, 10.103.1.0/2425s

usr0 — User/client tunnel (port 51821, NM-managed)

Local: 172.16.100.1/24

PeerIPName
1172.16.100.2trufa-lhe
2172.16.100.3bluelily-lhe
3172.16.100.4(unnamed)
4172.16.100.5dan1
5172.16.100.6dan2
6172.16.100.7dherbolt-mac

wg0 — Site-to-site tunnel (port 51820, wg-quick, boot-enabled)

Local: 172.168.200.1/24 — routes 172.168.10.0/24, 172.168.11.0/24, 172.168.100.0/24

Firewall

Zones

ZoneInterfacesTargetNotes
trustedenp3s0, wg0ACCEPTmasquerade enabled
wireguard-srvsrv0ACCEPTserver peers
wireguard-usrusr0ACCEPTuser peers, masquerade
block(default, none)REJECT-

Policies

  • wireguard-srv-in/out: bidirectional ACCEPT between ANY and wireguard-srv
  • wireguard-user-out: ACCEPT from wireguard-usr to ANY with masquerade
  • allow-host-ipv6: neighbor discovery / router advertisement

Key Services

ServicePurpose
sshdSSH
namedBIND DNS
wg-quick@wg0WireGuard tunnel
NetworkManagerManages srv0, usr0
firewalldnftables backend
chronydNTP
sssdAuth (Kerberos)
auditdAudit logging
pmcd/pmie/pmloggerPCP monitoring
SELinuxEnforcing
Failed unitsNone

Issues

  1. Fedora 42 EOL — no security updates since 2026-05-27

    dnf install -y dnf-plugin-system-upgrade
    dnf system-upgrade download --releasever=43
    dnf system-upgrade reboot
    
  2. wg0 dual management — autoconnect=false in NM but wg-quick@wg0 enabled at boot

    # Option A: let wg-quick manage wg0 (current), hide from NM
    echo -e "[keyfile]\nunmanaged-devices=interface-name:wg0" > /etc/NetworkManager/conf.d/unmanaged-wg0.conf
    nmcli general reload
    # Option B: migrate to NM-managed, disable wg-quick
    systemctl disable --now wg-quick@wg0
    nmcli con import type wireguard file /etc/wireguard/wg0.conf
    

Security Notes

  • Trusted zone on enp3s0 — all traffic accepted on physical interface

    # Create restrictive zone for VPN gateway
    firewall-cmd --permanent --new-zone=vpn-gw
    firewall-cmd --permanent --zone=vpn-gw --add-service=ssh
    firewall-cmd --permanent --zone=vpn-gw --add-service=dns
    firewall-cmd --permanent --zone=vpn-gw --add-port={5252/udp,51820/udp,51821/udp}
    firewall-cmd --permanent --zone=trusted --remove-interface=enp3s0
    firewall-cmd --permanent --zone=vpn-gw --add-interface=enp3s0
    firewall-cmd --reload
    
  • Password auth enabled in SSH — commented out, defaults to yes

    echo "PasswordAuthentication no" > /etc/ssh/sshd_config.d/60-no-password.conf
    systemctl reload sshd
    
  • WireGuard private key permissions

    chmod 0600 /etc/NetworkManager/system-connections/srv0.nmconnection
    chmod 0600 /etc/NetworkManager/system-connections/usr0.nmconnection
    chmod 0600 /etc/wireguard/wg0.conf
    
  • BIND recursive resolver — verify not open to untrusted networks

    # Check listen addresses and recursion ACL
    grep -E '(listen-on|allow-recursion|allow-query)' /etc/named.conf
    # Should be restricted to localhost and bridge subnet:
    # listen-on { 127.0.0.1; 172.16.31.30; };
    # allow-recursion { 127.0.0.0/8; 172.16.31.0/24; 172.16.40.0/24; 172.16.100.0/24; };
    
  • wireguard-usr zone — all user peers get unrestricted access

    # If user peers should only reach specific subnets, replace ACCEPT with rules:
    firewall-cmd --permanent --zone=wireguard-usr --set-target=DROP
    firewall-cmd --permanent --zone=wireguard-usr --add-rich-rule='rule family="ipv4" destination address="172.16.31.0/24" accept'
    firewall-cmd --permanent --zone=wireguard-usr --add-service={dns,ssh}
    firewall-cmd --reload
    
  • No fail2ban

    dnf install -y fail2ban
    cat > /etc/fail2ban/jail.local << 'EOF'
    [sshd]
    enabled = true
    maxretry = 5
    bantime = 3600
    EOF
    systemctl enable --now fail2ban
    

Performance Notes

  • No tuned profile active — set tuned-adm profile network-latency for VPN gateway (reduces latency jitter, optimizes network stack parameters)
  • 4 vCPUs — adequate for WireGuard + BIND at current peer count (6 user peers, 2 server peers). WireGuard is kernel-space and efficient
  • BIND query caching — verify max-cache-size is tuned for available RAM. Default can consume up to 90% of available memory
  • conntrack table — VPN gateway with masquerade needs adequate conntrack table size. Check net.netfilter.nf_conntrack_max for high peer/connection counts

Sosreport Plugins

Enabled extras: named, sssd, tuned, conntrack

All recommended plugins for this server role are already enabled. No sos plugins exist for WireGuard or nftables — WireGuard config collected via networking/networkmanager (enabled by default), nftables rules via firewall_tables/firewalld (enabled by default).

sos report -e named,sssd,tuned,conntrack