Files
madomeda/RULES_GUIDE.md
T
2025-10-23 20:14:21 +02:00

11 KiB

Writing Custom Rules

Overview

Madomeda uses a self-contained, JSON-based rules system. Rules are processed in priority order and can normalize, validate, and transform frontmatter fields without any code changes.

Rule File Location

Rules are stored in: ~/.config/madomeda/rules/

Each rule is a separate JSON file. The filename is used for sorting (alphabetically), but the priority field determines execution order.

Rule Structure

{
  "name": "rule_name",
  "description": "Human-readable description",
  "field": "field_name",
  "priority": 10,
  "action": "action_type",
  "pattern": "regex_pattern",
  "replacement": "replacement_string",
  "llm_prompt": "Prompt for LLM if heuristics fail"
}

Required Fields

  • name: Unique identifier for the rule
  • description: What the rule does
  • field: Which frontmatter field this applies to (use "*" for all fields)
  • priority: Execution order (lower number = earlier execution, typically 1-100)
  • action: What the rule does (see Actions below)

Optional Fields

  • pattern: Regex pattern for matching/replacing
  • replacement: Replacement string (can use capture groups like \g<0>, \1, etc.)
  • transform: Transformation to apply (lower, upper)
  • multiline: Whether regex uses multiline mode (default: false)
  • split_on: Character to split on (e.g., "," for comma-separated values)
  • from: Source field name (for rename actions)
  • to: Target field name (for rename actions)
  • llm_prompt: Text to send to LLM if this rule's value needs inference

Actions

1. normalize_keys

Apply transformation to all frontmatter keys.

Use case: Ensure all keys are lowercase.

Example:

{
  "name": "lowercase_keys",
  "description": "Convert all keys to lowercase",
  "field": "*",
  "priority": 1,
  "action": "normalize_keys",
  "transform": "lower",
  "llm_prompt": "Ensure all keys are lowercase"
}

Alternative with regex:

{
  "name": "remove_spaces_from_keys",
  "description": "Remove spaces from keys",
  "field": "*",
  "priority": 2,
  "action": "normalize_keys",
  "pattern": " ",
  "replacement": "_"
}

2. rename_field

Rename one field to another.

Use case: Standardize field names (e.g., tagtags, summarydescription).

Example:

{
  "name": "tag_to_tags",
  "description": "Rename 'tag' to 'tags'",
  "field": "tag",
  "priority": 2,
  "action": "rename_field",
  "from": "tag",
  "to": "tags",
  "llm_prompt": "Convert tag field to tags array"
}

Note: Rename rules execute before other field-specific rules, so subsequent rules can process the renamed field.

3. normalize_value

Transform field values using regex, transforms, or splitting.

Use case: Standardize tag format, clean up values, convert formats.

Example 1: Split comma-separated to list

{
  "name": "tags_split",
  "description": "Convert comma-separated tags to list",
  "field": "tags",
  "priority": 10,
  "action": "normalize_value",
  "split_on": ",",
  "pattern": ".*",
  "replacement": "\\g<0>"
}

Example 2: Replace spaces with underscores

{
  "name": "tags_no_spaces",
  "description": "Replace spaces with underscores in tags",
  "field": "tags",
  "priority": 20,
  "action": "normalize_value",
  "pattern": " ",
  "replacement": "_"
}

Example 3: Lowercase transformation

{
  "name": "tags_lowercase",
  "description": "Convert tags to lowercase",
  "field": "tags",
  "priority": 30,
  "action": "normalize_value",
  "transform": "lower"
}

Example 4: Remove invalid characters

{
  "name": "tags_alphanumeric_only",
  "description": "Keep only alphanumeric and underscores",
  "field": "tags",
  "priority": 40,
  "action": "normalize_value",
  "pattern": "[^a-z0-9_]",
  "replacement": ""
}

4. validate

Check if field values match a pattern.

Use case: Ensure values conform to expected format.

Example:

{
  "name": "tags_validate",
  "description": "Validate tag format",
  "field": "tags",
  "priority": 50,
  "action": "validate",
  "pattern": "^[a-z0-9_]+$",
  "multiline": false
}

Note: Validation rules mark frontmatter as non-conformant if they fail, but don't modify values.

Priority System

Rules execute in priority order (lowest number first):

  • 1-9: Global transformations (keys, field renames)
  • 10-19: Format conversions (split lists, type changes)
  • 20-29: Character replacements (spaces, hyphens)
  • 30-39: Case transformations
  • 40-49: Character removal (invalid chars)
  • 50-99: Validation

Recommended naming convention:

01_lowercase_keys.json       # Priority 1
10_tags_format_list.json     # Priority 10
20_tags_replace_spaces.json  # Priority 20
50_tags_validate.json        # Priority 50

Processing Flow

  1. Global rules (field: "*") execute first
  2. Rename rules execute second (creates new fields)
  3. Other rules execute in priority order per field

Complete Example: Tag Normalization Chain

Here's how multiple rules work together to normalize tags:

Input:

Tag: Machine-Learning, Deep Learning, AI/ML

Rules (in priority order):

// 02_tag_to_tags.json
{
  "name": "tag_to_tags",
  "field": "tag",
  "priority": 2,
  "action": "rename_field",
  "from": "tag",
  "to": "tags"
}

After this rule: tags: "Machine-Learning, Deep Learning, AI/ML"

// 10_tags_format_list.json
{
  "name": "tags_split",
  "field": "tags",
  "priority": 10,
  "action": "normalize_value",
  "split_on": ","
}

After this rule: tags: ["Machine-Learning", "Deep Learning", "AI/ML"]

// 20_tags_replace_spaces.json
{
  "name": "tags_spaces",
  "field": "tags",
  "priority": 20,
  "action": "normalize_value",
  "pattern": " ",
  "replacement": "_"
}

After this rule: tags: ["Machine-Learning", "Deep_Learning", "AI/ML"]

// 21_tags_replace_hyphens.json
{
  "name": "tags_hyphens",
  "field": "tags",
  "priority": 21,
  "action": "normalize_value",
  "pattern": "-",
  "replacement": "_"
}

After this rule: tags: ["Machine_Learning", "Deep_Learning", "AI/ML"]

// 30_tags_lowercase.json
{
  "name": "tags_lowercase",
  "field": "tags",
  "priority": 30,
  "action": "normalize_value",
  "transform": "lower"
}

After this rule: tags: ["machine_learning", "deep_learning", "ai/ml"]

// 40_tags_remove_invalid.json
{
  "name": "tags_clean",
  "field": "tags",
  "priority": 40,
  "action": "normalize_value",
  "pattern": "[^a-z0-9_]",
  "replacement": ""
}

Final result: tags: ["machine_learning", "deep_learning", "aiml"]

Field Types

Rules apply to different value types:

String Fields

Rules apply to the whole string:

{
  "field": "title",
  "action": "normalize_value",
  "transform": "lower"
}

List Fields

Rules apply to each item in the list:

{
  "field": "keywords",
  "action": "normalize_value",
  "pattern": " ",
  "replacement": "_"
}

Each keyword gets spaces replaced with underscores.

Converting String to List

Use split_on:

{
  "field": "authors",
  "action": "normalize_value",
  "split_on": ";",
  "pattern": ".*",
  "replacement": "\\g<0>"
}

Input: "John Doe; Jane Smith" Output: ["John Doe", "Jane Smith"]

Regular Expression Tips

Capture Groups

{
  "pattern": "(\\d+)",
  "replacement": "v\\1"
}

Input: "123" Output: "v123"

Case-Insensitive Matching

Use inline flag:

{
  "pattern": "(?i)todo",
  "replacement": "TODO"
}

Match Whole String

{
  "pattern": "^[a-z]+$"
}

Remove Leading/Trailing Whitespace

{
  "pattern": "^\\s+|\\s+$",
  "replacement": ""
}

LLM Integration

The llm_prompt field provides guidance when the LLM needs to generate or fix values:

{
  "name": "tags_validate",
  "field": "tags",
  "action": "validate",
  "pattern": "^[a-z0-9_]+$",
  "llm_prompt": "Generate tags using only lowercase letters, numbers, and underscores. Prefer selecting from existing repository tags when appropriate."
}

When using ai strategy in templates, the LLM receives:

  • The rule's llm_prompt
  • List of existing tags in repository
  • Document content
  • Structured output schema

Testing Rules

Test with --whatif

madomeda --whatif

Shows what would change without modifying files.

Test Specific Rule

Create a Python script:

from core.rules_processor import RulesProcessor
from pathlib import Path

rp = RulesProcessor(Path.home() / '.config' / 'madomeda' / 'rules')

test_data = {
    'Title': 'My Document',
    'Tag': 'Python, Machine-Learning'
}

result, conformant, violations = rp.apply_rules(test_data)
print(f"Result: {result}")
print(f"Conformant: {conformant}")
print(f"Violations: {violations}")

Common Patterns

Email Validation

{
  "field": "author_email",
  "action": "validate",
  "pattern": "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$"
}

Date Format Validation (YYYY-MM-DD)

{
  "field": "date",
  "action": "validate",
  "pattern": "^\\d{4}-\\d{2}-\\d{2}$"
}

URL Scheme Enforcement

{
  "field": "url",
  "action": "normalize_value",
  "pattern": "^(?!https?://)(.*)",
  "replacement": "https://\\1"
}

Remove HTML Tags

{
  "field": "description",
  "action": "normalize_value",
  "pattern": "<[^>]+>",
  "replacement": ""
}

Advanced: Custom Field Rules

Ensure Boolean String Format

{
  "field": "published",
  "action": "normalize_value",
  "pattern": "^(true|false|yes|no|1|0)$",
  "replacement": "\\1",
  "transform": "lower"
}

Then normalize:

{
  "field": "published",
  "priority": 31,
  "action": "normalize_value",
  "pattern": "yes|1|true",
  "replacement": "true"
}

Slug Generation

{
  "field": "slug",
  "action": "normalize_value",
  "pattern": "[^a-z0-9-]",
  "replacement": "",
  "transform": "lower"
}

Troubleshooting

Rule Not Executing

  • Check field matches exact field name
  • Verify priority is in expected range
  • Ensure JSON is valid (use jsonlint)

Wrong Execution Order

  • Lower priority number executes first
  • Rename rules always execute before other rules for same field
  • Global rules (field: "*") execute before field-specific

Regex Not Matching

  • Test regex at https://regex101.com/
  • Escape special characters: . * + ? ^ $ { } ( ) | [ ] \
  • Use multiline: true for multiline patterns

Value Not Changing

  • Check if earlier rule already modified it
  • Verify pattern actually matches the value
  • Ensure replacement is specified for normalize_value

Best Practices

  1. Use priority ranges - Leave gaps (10, 20, 30) to insert rules later
  2. Name files by priority - 01_, 10_, 20_ for easy sorting
  3. Test incrementally - Add one rule at a time
  4. Document complex regex - Use description field
  5. Provide LLM prompts - Help AI understand the rule intent
  6. Validate after normalize - Use separate validate rule at higher priority

See Also