563 lines
11 KiB
Markdown
563 lines
11 KiB
Markdown
|
|
# Writing Custom Rules
|
||
|
|
|
||
|
|
## Overview
|
||
|
|
|
||
|
|
Madomeda uses a self-contained, JSON-based rules system. Rules are processed in priority order and can normalize, validate, and transform frontmatter fields without any code changes.
|
||
|
|
|
||
|
|
## Rule File Location
|
||
|
|
|
||
|
|
Rules are stored in: `~/.config/madomeda/rules/`
|
||
|
|
|
||
|
|
Each rule is a separate JSON file. The filename is used for sorting (alphabetically), but the `priority` field determines execution order.
|
||
|
|
|
||
|
|
## Rule Structure
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "rule_name",
|
||
|
|
"description": "Human-readable description",
|
||
|
|
"field": "field_name",
|
||
|
|
"priority": 10,
|
||
|
|
"action": "action_type",
|
||
|
|
"pattern": "regex_pattern",
|
||
|
|
"replacement": "replacement_string",
|
||
|
|
"llm_prompt": "Prompt for LLM if heuristics fail"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Required Fields
|
||
|
|
|
||
|
|
- **name**: Unique identifier for the rule
|
||
|
|
- **description**: What the rule does
|
||
|
|
- **field**: Which frontmatter field this applies to (use `"*"` for all fields)
|
||
|
|
- **priority**: Execution order (lower number = earlier execution, typically 1-100)
|
||
|
|
- **action**: What the rule does (see Actions below)
|
||
|
|
|
||
|
|
### Optional Fields
|
||
|
|
|
||
|
|
- **pattern**: Regex pattern for matching/replacing
|
||
|
|
- **replacement**: Replacement string (can use capture groups like `\g<0>`, `\1`, etc.)
|
||
|
|
- **transform**: Transformation to apply (`lower`, `upper`)
|
||
|
|
- **multiline**: Whether regex uses multiline mode (default: false)
|
||
|
|
- **split_on**: Character to split on (e.g., `","` for comma-separated values)
|
||
|
|
- **from**: Source field name (for rename actions)
|
||
|
|
- **to**: Target field name (for rename actions)
|
||
|
|
- **llm_prompt**: Text to send to LLM if this rule's value needs inference
|
||
|
|
|
||
|
|
## Actions
|
||
|
|
|
||
|
|
### 1. normalize_keys
|
||
|
|
|
||
|
|
Apply transformation to all frontmatter keys.
|
||
|
|
|
||
|
|
**Use case**: Ensure all keys are lowercase.
|
||
|
|
|
||
|
|
**Example:**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "lowercase_keys",
|
||
|
|
"description": "Convert all keys to lowercase",
|
||
|
|
"field": "*",
|
||
|
|
"priority": 1,
|
||
|
|
"action": "normalize_keys",
|
||
|
|
"transform": "lower",
|
||
|
|
"llm_prompt": "Ensure all keys are lowercase"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Alternative with regex:**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "remove_spaces_from_keys",
|
||
|
|
"description": "Remove spaces from keys",
|
||
|
|
"field": "*",
|
||
|
|
"priority": 2,
|
||
|
|
"action": "normalize_keys",
|
||
|
|
"pattern": " ",
|
||
|
|
"replacement": "_"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### 2. rename_field
|
||
|
|
|
||
|
|
Rename one field to another.
|
||
|
|
|
||
|
|
**Use case**: Standardize field names (e.g., `tag` → `tags`, `summary` → `description`).
|
||
|
|
|
||
|
|
**Example:**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tag_to_tags",
|
||
|
|
"description": "Rename 'tag' to 'tags'",
|
||
|
|
"field": "tag",
|
||
|
|
"priority": 2,
|
||
|
|
"action": "rename_field",
|
||
|
|
"from": "tag",
|
||
|
|
"to": "tags",
|
||
|
|
"llm_prompt": "Convert tag field to tags array"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Note**: Rename rules execute before other field-specific rules, so subsequent rules can process the renamed field.
|
||
|
|
|
||
|
|
### 3. normalize_value
|
||
|
|
|
||
|
|
Transform field values using regex, transforms, or splitting.
|
||
|
|
|
||
|
|
**Use case**: Standardize tag format, clean up values, convert formats.
|
||
|
|
|
||
|
|
**Example 1: Split comma-separated to list**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_split",
|
||
|
|
"description": "Convert comma-separated tags to list",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 10,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"split_on": ",",
|
||
|
|
"pattern": ".*",
|
||
|
|
"replacement": "\\g<0>"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Example 2: Replace spaces with underscores**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_no_spaces",
|
||
|
|
"description": "Replace spaces with underscores in tags",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 20,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": " ",
|
||
|
|
"replacement": "_"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Example 3: Lowercase transformation**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_lowercase",
|
||
|
|
"description": "Convert tags to lowercase",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 30,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"transform": "lower"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Example 4: Remove invalid characters**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_alphanumeric_only",
|
||
|
|
"description": "Keep only alphanumeric and underscores",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 40,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "[^a-z0-9_]",
|
||
|
|
"replacement": ""
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### 4. validate
|
||
|
|
|
||
|
|
Check if field values match a pattern.
|
||
|
|
|
||
|
|
**Use case**: Ensure values conform to expected format.
|
||
|
|
|
||
|
|
**Example:**
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_validate",
|
||
|
|
"description": "Validate tag format",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 50,
|
||
|
|
"action": "validate",
|
||
|
|
"pattern": "^[a-z0-9_]+$",
|
||
|
|
"multiline": false
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Note**: Validation rules mark frontmatter as non-conformant if they fail, but don't modify values.
|
||
|
|
|
||
|
|
## Priority System
|
||
|
|
|
||
|
|
Rules execute in priority order (lowest number first):
|
||
|
|
|
||
|
|
- **1-9**: Global transformations (keys, field renames)
|
||
|
|
- **10-19**: Format conversions (split lists, type changes)
|
||
|
|
- **20-29**: Character replacements (spaces, hyphens)
|
||
|
|
- **30-39**: Case transformations
|
||
|
|
- **40-49**: Character removal (invalid chars)
|
||
|
|
- **50-99**: Validation
|
||
|
|
|
||
|
|
**Recommended naming convention:**
|
||
|
|
```
|
||
|
|
01_lowercase_keys.json # Priority 1
|
||
|
|
10_tags_format_list.json # Priority 10
|
||
|
|
20_tags_replace_spaces.json # Priority 20
|
||
|
|
50_tags_validate.json # Priority 50
|
||
|
|
```
|
||
|
|
|
||
|
|
## Processing Flow
|
||
|
|
|
||
|
|
1. **Global rules** (`field: "*"`) execute first
|
||
|
|
2. **Rename rules** execute second (creates new fields)
|
||
|
|
3. **Other rules** execute in priority order per field
|
||
|
|
|
||
|
|
## Complete Example: Tag Normalization Chain
|
||
|
|
|
||
|
|
Here's how multiple rules work together to normalize tags:
|
||
|
|
|
||
|
|
**Input:**
|
||
|
|
```yaml
|
||
|
|
Tag: Machine-Learning, Deep Learning, AI/ML
|
||
|
|
```
|
||
|
|
|
||
|
|
**Rules (in priority order):**
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 02_tag_to_tags.json
|
||
|
|
{
|
||
|
|
"name": "tag_to_tags",
|
||
|
|
"field": "tag",
|
||
|
|
"priority": 2,
|
||
|
|
"action": "rename_field",
|
||
|
|
"from": "tag",
|
||
|
|
"to": "tags"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
After this rule: `tags: "Machine-Learning, Deep Learning, AI/ML"`
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 10_tags_format_list.json
|
||
|
|
{
|
||
|
|
"name": "tags_split",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 10,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"split_on": ","
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
After this rule: `tags: ["Machine-Learning", "Deep Learning", "AI/ML"]`
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 20_tags_replace_spaces.json
|
||
|
|
{
|
||
|
|
"name": "tags_spaces",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 20,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": " ",
|
||
|
|
"replacement": "_"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
After this rule: `tags: ["Machine-Learning", "Deep_Learning", "AI/ML"]`
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 21_tags_replace_hyphens.json
|
||
|
|
{
|
||
|
|
"name": "tags_hyphens",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 21,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "-",
|
||
|
|
"replacement": "_"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
After this rule: `tags: ["Machine_Learning", "Deep_Learning", "AI/ML"]`
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 30_tags_lowercase.json
|
||
|
|
{
|
||
|
|
"name": "tags_lowercase",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 30,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"transform": "lower"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
After this rule: `tags: ["machine_learning", "deep_learning", "ai/ml"]`
|
||
|
|
|
||
|
|
```json
|
||
|
|
// 40_tags_remove_invalid.json
|
||
|
|
{
|
||
|
|
"name": "tags_clean",
|
||
|
|
"field": "tags",
|
||
|
|
"priority": 40,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "[^a-z0-9_]",
|
||
|
|
"replacement": ""
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Final result:** `tags: ["machine_learning", "deep_learning", "aiml"]`
|
||
|
|
|
||
|
|
## Field Types
|
||
|
|
|
||
|
|
Rules apply to different value types:
|
||
|
|
|
||
|
|
### String Fields
|
||
|
|
|
||
|
|
Rules apply to the whole string:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "title",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"transform": "lower"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### List Fields
|
||
|
|
|
||
|
|
Rules apply to each item in the list:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "keywords",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": " ",
|
||
|
|
"replacement": "_"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
Each keyword gets spaces replaced with underscores.
|
||
|
|
|
||
|
|
### Converting String to List
|
||
|
|
|
||
|
|
Use `split_on`:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "authors",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"split_on": ";",
|
||
|
|
"pattern": ".*",
|
||
|
|
"replacement": "\\g<0>"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
Input: `"John Doe; Jane Smith"`
|
||
|
|
Output: `["John Doe", "Jane Smith"]`
|
||
|
|
|
||
|
|
## Regular Expression Tips
|
||
|
|
|
||
|
|
### Capture Groups
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"pattern": "(\\d+)",
|
||
|
|
"replacement": "v\\1"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
Input: `"123"`
|
||
|
|
Output: `"v123"`
|
||
|
|
|
||
|
|
### Case-Insensitive Matching
|
||
|
|
|
||
|
|
Use inline flag:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"pattern": "(?i)todo",
|
||
|
|
"replacement": "TODO"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Match Whole String
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"pattern": "^[a-z]+$"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Remove Leading/Trailing Whitespace
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"pattern": "^\\s+|\\s+$",
|
||
|
|
"replacement": ""
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
## LLM Integration
|
||
|
|
|
||
|
|
The `llm_prompt` field provides guidance when the LLM needs to generate or fix values:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "tags_validate",
|
||
|
|
"field": "tags",
|
||
|
|
"action": "validate",
|
||
|
|
"pattern": "^[a-z0-9_]+$",
|
||
|
|
"llm_prompt": "Generate tags using only lowercase letters, numbers, and underscores. Prefer selecting from existing repository tags when appropriate."
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
When using `ai` strategy in templates, the LLM receives:
|
||
|
|
- The rule's `llm_prompt`
|
||
|
|
- List of existing tags in repository
|
||
|
|
- Document content
|
||
|
|
- Structured output schema
|
||
|
|
|
||
|
|
## Testing Rules
|
||
|
|
|
||
|
|
### Test with --whatif
|
||
|
|
|
||
|
|
```bash
|
||
|
|
madomeda --whatif
|
||
|
|
```
|
||
|
|
|
||
|
|
Shows what would change without modifying files.
|
||
|
|
|
||
|
|
### Test Specific Rule
|
||
|
|
|
||
|
|
Create a Python script:
|
||
|
|
|
||
|
|
```python
|
||
|
|
from core.rules_processor import RulesProcessor
|
||
|
|
from pathlib import Path
|
||
|
|
|
||
|
|
rp = RulesProcessor(Path.home() / '.config' / 'madomeda' / 'rules')
|
||
|
|
|
||
|
|
test_data = {
|
||
|
|
'Title': 'My Document',
|
||
|
|
'Tag': 'Python, Machine-Learning'
|
||
|
|
}
|
||
|
|
|
||
|
|
result, conformant, violations = rp.apply_rules(test_data)
|
||
|
|
print(f"Result: {result}")
|
||
|
|
print(f"Conformant: {conformant}")
|
||
|
|
print(f"Violations: {violations}")
|
||
|
|
```
|
||
|
|
|
||
|
|
## Common Patterns
|
||
|
|
|
||
|
|
### Email Validation
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "author_email",
|
||
|
|
"action": "validate",
|
||
|
|
"pattern": "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Date Format Validation (YYYY-MM-DD)
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "date",
|
||
|
|
"action": "validate",
|
||
|
|
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### URL Scheme Enforcement
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "url",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "^(?!https?://)(.*)",
|
||
|
|
"replacement": "https://\\1"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Remove HTML Tags
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "description",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "<[^>]+>",
|
||
|
|
"replacement": ""
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
## Advanced: Custom Field Rules
|
||
|
|
|
||
|
|
### Ensure Boolean String Format
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "published",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "^(true|false|yes|no|1|0)$",
|
||
|
|
"replacement": "\\1",
|
||
|
|
"transform": "lower"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
Then normalize:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "published",
|
||
|
|
"priority": 31,
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "yes|1|true",
|
||
|
|
"replacement": "true"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
### Slug Generation
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"field": "slug",
|
||
|
|
"action": "normalize_value",
|
||
|
|
"pattern": "[^a-z0-9-]",
|
||
|
|
"replacement": "",
|
||
|
|
"transform": "lower"
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
## Troubleshooting
|
||
|
|
|
||
|
|
### Rule Not Executing
|
||
|
|
|
||
|
|
- Check `field` matches exact field name
|
||
|
|
- Verify `priority` is in expected range
|
||
|
|
- Ensure JSON is valid (use `jsonlint`)
|
||
|
|
|
||
|
|
### Wrong Execution Order
|
||
|
|
|
||
|
|
- Lower priority number executes first
|
||
|
|
- Rename rules always execute before other rules for same field
|
||
|
|
- Global rules (`field: "*"`) execute before field-specific
|
||
|
|
|
||
|
|
### Regex Not Matching
|
||
|
|
|
||
|
|
- Test regex at https://regex101.com/
|
||
|
|
- Escape special characters: `. * + ? ^ $ { } ( ) | [ ] \`
|
||
|
|
- Use `multiline: true` for multiline patterns
|
||
|
|
|
||
|
|
### Value Not Changing
|
||
|
|
|
||
|
|
- Check if earlier rule already modified it
|
||
|
|
- Verify pattern actually matches the value
|
||
|
|
- Ensure `replacement` is specified for `normalize_value`
|
||
|
|
|
||
|
|
## Best Practices
|
||
|
|
|
||
|
|
1. **Use priority ranges** - Leave gaps (10, 20, 30) to insert rules later
|
||
|
|
2. **Name files by priority** - `01_`, `10_`, `20_` for easy sorting
|
||
|
|
3. **Test incrementally** - Add one rule at a time
|
||
|
|
4. **Document complex regex** - Use `description` field
|
||
|
|
5. **Provide LLM prompts** - Help AI understand the rule intent
|
||
|
|
6. **Validate after normalize** - Use separate validate rule at higher priority
|
||
|
|
|
||
|
|
## See Also
|
||
|
|
|
||
|
|
- [CONFIGURATION.md](CONFIGURATION.md) - Configuration directory
|
||
|
|
- [STRUCTURE.md](STRUCTURE.md) - How rules fit into the system
|
||
|
|
- [USAGE.md](USAGE.md) - Using rules with madomeda
|