11 KiB
Writing Custom Rules
Overview
Madomeda uses a self-contained, JSON-based rules system. Rules are processed in priority order and can normalize, validate, and transform frontmatter fields without any code changes.
Rule File Location
Rules are stored in: ~/.config/madomeda/rules/
Each rule is a separate JSON file. The filename is used for sorting (alphabetically), but the priority field determines execution order.
Rule Structure
{
"name": "rule_name",
"description": "Human-readable description",
"field": "field_name",
"priority": 10,
"action": "action_type",
"pattern": "regex_pattern",
"replacement": "replacement_string",
"llm_prompt": "Prompt for LLM if heuristics fail"
}
Required Fields
- name: Unique identifier for the rule
- description: What the rule does
- field: Which frontmatter field this applies to (use
"*"for all fields) - priority: Execution order (lower number = earlier execution, typically 1-100)
- action: What the rule does (see Actions below)
Optional Fields
- pattern: Regex pattern for matching/replacing
- replacement: Replacement string (can use capture groups like
\g<0>,\1, etc.) - transform: Transformation to apply (
lower,upper) - multiline: Whether regex uses multiline mode (default: false)
- split_on: Character to split on (e.g.,
","for comma-separated values) - from: Source field name (for rename actions)
- to: Target field name (for rename actions)
- llm_prompt: Text to send to LLM if this rule's value needs inference
Actions
1. normalize_keys
Apply transformation to all frontmatter keys.
Use case: Ensure all keys are lowercase.
Example:
{
"name": "lowercase_keys",
"description": "Convert all keys to lowercase",
"field": "*",
"priority": 1,
"action": "normalize_keys",
"transform": "lower",
"llm_prompt": "Ensure all keys are lowercase"
}
Alternative with regex:
{
"name": "remove_spaces_from_keys",
"description": "Remove spaces from keys",
"field": "*",
"priority": 2,
"action": "normalize_keys",
"pattern": " ",
"replacement": "_"
}
2. rename_field
Rename one field to another.
Use case: Standardize field names (e.g., tag → tags, summary → description).
Example:
{
"name": "tag_to_tags",
"description": "Rename 'tag' to 'tags'",
"field": "tag",
"priority": 2,
"action": "rename_field",
"from": "tag",
"to": "tags",
"llm_prompt": "Convert tag field to tags array"
}
Note: Rename rules execute before other field-specific rules, so subsequent rules can process the renamed field.
3. normalize_value
Transform field values using regex, transforms, or splitting.
Use case: Standardize tag format, clean up values, convert formats.
Example 1: Split comma-separated to list
{
"name": "tags_split",
"description": "Convert comma-separated tags to list",
"field": "tags",
"priority": 10,
"action": "normalize_value",
"split_on": ",",
"pattern": ".*",
"replacement": "\\g<0>"
}
Example 2: Replace spaces with underscores
{
"name": "tags_no_spaces",
"description": "Replace spaces with underscores in tags",
"field": "tags",
"priority": 20,
"action": "normalize_value",
"pattern": " ",
"replacement": "_"
}
Example 3: Lowercase transformation
{
"name": "tags_lowercase",
"description": "Convert tags to lowercase",
"field": "tags",
"priority": 30,
"action": "normalize_value",
"transform": "lower"
}
Example 4: Remove invalid characters
{
"name": "tags_alphanumeric_only",
"description": "Keep only alphanumeric and underscores",
"field": "tags",
"priority": 40,
"action": "normalize_value",
"pattern": "[^a-z0-9_]",
"replacement": ""
}
4. validate
Check if field values match a pattern.
Use case: Ensure values conform to expected format.
Example:
{
"name": "tags_validate",
"description": "Validate tag format",
"field": "tags",
"priority": 50,
"action": "validate",
"pattern": "^[a-z0-9_]+$",
"multiline": false
}
Note: Validation rules mark frontmatter as non-conformant if they fail, but don't modify values.
Priority System
Rules execute in priority order (lowest number first):
- 1-9: Global transformations (keys, field renames)
- 10-19: Format conversions (split lists, type changes)
- 20-29: Character replacements (spaces, hyphens)
- 30-39: Case transformations
- 40-49: Character removal (invalid chars)
- 50-99: Validation
Recommended naming convention:
01_lowercase_keys.json # Priority 1
10_tags_format_list.json # Priority 10
20_tags_replace_spaces.json # Priority 20
50_tags_validate.json # Priority 50
Processing Flow
- Global rules (
field: "*") execute first - Rename rules execute second (creates new fields)
- Other rules execute in priority order per field
Complete Example: Tag Normalization Chain
Here's how multiple rules work together to normalize tags:
Input:
Tag: Machine-Learning, Deep Learning, AI/ML
Rules (in priority order):
// 02_tag_to_tags.json
{
"name": "tag_to_tags",
"field": "tag",
"priority": 2,
"action": "rename_field",
"from": "tag",
"to": "tags"
}
After this rule: tags: "Machine-Learning, Deep Learning, AI/ML"
// 10_tags_format_list.json
{
"name": "tags_split",
"field": "tags",
"priority": 10,
"action": "normalize_value",
"split_on": ","
}
After this rule: tags: ["Machine-Learning", "Deep Learning", "AI/ML"]
// 20_tags_replace_spaces.json
{
"name": "tags_spaces",
"field": "tags",
"priority": 20,
"action": "normalize_value",
"pattern": " ",
"replacement": "_"
}
After this rule: tags: ["Machine-Learning", "Deep_Learning", "AI/ML"]
// 21_tags_replace_hyphens.json
{
"name": "tags_hyphens",
"field": "tags",
"priority": 21,
"action": "normalize_value",
"pattern": "-",
"replacement": "_"
}
After this rule: tags: ["Machine_Learning", "Deep_Learning", "AI/ML"]
// 30_tags_lowercase.json
{
"name": "tags_lowercase",
"field": "tags",
"priority": 30,
"action": "normalize_value",
"transform": "lower"
}
After this rule: tags: ["machine_learning", "deep_learning", "ai/ml"]
// 40_tags_remove_invalid.json
{
"name": "tags_clean",
"field": "tags",
"priority": 40,
"action": "normalize_value",
"pattern": "[^a-z0-9_]",
"replacement": ""
}
Final result: tags: ["machine_learning", "deep_learning", "aiml"]
Field Types
Rules apply to different value types:
String Fields
Rules apply to the whole string:
{
"field": "title",
"action": "normalize_value",
"transform": "lower"
}
List Fields
Rules apply to each item in the list:
{
"field": "keywords",
"action": "normalize_value",
"pattern": " ",
"replacement": "_"
}
Each keyword gets spaces replaced with underscores.
Converting String to List
Use split_on:
{
"field": "authors",
"action": "normalize_value",
"split_on": ";",
"pattern": ".*",
"replacement": "\\g<0>"
}
Input: "John Doe; Jane Smith"
Output: ["John Doe", "Jane Smith"]
Regular Expression Tips
Capture Groups
{
"pattern": "(\\d+)",
"replacement": "v\\1"
}
Input: "123"
Output: "v123"
Case-Insensitive Matching
Use inline flag:
{
"pattern": "(?i)todo",
"replacement": "TODO"
}
Match Whole String
{
"pattern": "^[a-z]+$"
}
Remove Leading/Trailing Whitespace
{
"pattern": "^\\s+|\\s+$",
"replacement": ""
}
LLM Integration
The llm_prompt field provides guidance when the LLM needs to generate or fix values:
{
"name": "tags_validate",
"field": "tags",
"action": "validate",
"pattern": "^[a-z0-9_]+$",
"llm_prompt": "Generate tags using only lowercase letters, numbers, and underscores. Prefer selecting from existing repository tags when appropriate."
}
When using ai strategy in templates, the LLM receives:
- The rule's
llm_prompt - List of existing tags in repository
- Document content
- Structured output schema
Testing Rules
Test with --whatif
madomeda --whatif
Shows what would change without modifying files.
Test Specific Rule
Create a Python script:
from core.rules_processor import RulesProcessor
from pathlib import Path
rp = RulesProcessor(Path.home() / '.config' / 'madomeda' / 'rules')
test_data = {
'Title': 'My Document',
'Tag': 'Python, Machine-Learning'
}
result, conformant, violations = rp.apply_rules(test_data)
print(f"Result: {result}")
print(f"Conformant: {conformant}")
print(f"Violations: {violations}")
Common Patterns
Email Validation
{
"field": "author_email",
"action": "validate",
"pattern": "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$"
}
Date Format Validation (YYYY-MM-DD)
{
"field": "date",
"action": "validate",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
}
URL Scheme Enforcement
{
"field": "url",
"action": "normalize_value",
"pattern": "^(?!https?://)(.*)",
"replacement": "https://\\1"
}
Remove HTML Tags
{
"field": "description",
"action": "normalize_value",
"pattern": "<[^>]+>",
"replacement": ""
}
Advanced: Custom Field Rules
Ensure Boolean String Format
{
"field": "published",
"action": "normalize_value",
"pattern": "^(true|false|yes|no|1|0)$",
"replacement": "\\1",
"transform": "lower"
}
Then normalize:
{
"field": "published",
"priority": 31,
"action": "normalize_value",
"pattern": "yes|1|true",
"replacement": "true"
}
Slug Generation
{
"field": "slug",
"action": "normalize_value",
"pattern": "[^a-z0-9-]",
"replacement": "",
"transform": "lower"
}
Troubleshooting
Rule Not Executing
- Check
fieldmatches exact field name - Verify
priorityis in expected range - Ensure JSON is valid (use
jsonlint)
Wrong Execution Order
- Lower priority number executes first
- Rename rules always execute before other rules for same field
- Global rules (
field: "*") execute before field-specific
Regex Not Matching
- Test regex at https://regex101.com/
- Escape special characters:
. * + ? ^ $ { } ( ) | [ ] \ - Use
multiline: truefor multiline patterns
Value Not Changing
- Check if earlier rule already modified it
- Verify pattern actually matches the value
- Ensure
replacementis specified fornormalize_value
Best Practices
- Use priority ranges - Leave gaps (10, 20, 30) to insert rules later
- Name files by priority -
01_,10_,20_for easy sorting - Test incrementally - Add one rule at a time
- Document complex regex - Use
descriptionfield - Provide LLM prompts - Help AI understand the rule intent
- Validate after normalize - Use separate validate rule at higher priority
See Also
- CONFIGURATION.md - Configuration directory
- STRUCTURE.md - How rules fit into the system
- USAGE.md - Using rules with madomeda