Instrumentation & Metrics
Table of Contents
- Overview
- Quick Start
- Instrumentation Approaches
- Metrics Endpoints
- Scheduler Metrics
- Custom Metrics
- Configuration
- Integration with Prometheus & Grafana
- Complete Example
Overview
Mitsuki provides built-in instrumentation for monitoring application performance and behavior:
- HTTP requests: Request counts, latency percentiles, status codes
- Scheduled tasks: Execution counts, durations, failures
- Component calls: Method execution times, error rates
- System resources: CPU usage, memory consumption
- Custom metrics: Track application-specific operational data
Every metric is exposed at /metrics/prometheus in the Prometheus text format. /metrics returns a human-readable JSON summary of the built-in HTTP, component, scheduler and system metrics; custom metrics appear only in the Prometheus output.
Quick Start
1. Install
Instrumentation samples process CPU and memory through psutil, which ships as an optional extra:
pip install "mitsuki[metrics]"Enabling instrumentation without it fails at startup with an error naming the extra. Metrics endpoints, scheduler metrics and custom metrics work without it.
2. Enable in Configuration
Add to application.yml:
instrumentation:
enabled: true
track_memory: false
metrics:
enabled: true
path: /metrics
allowed_ips: [] # Empty list = allow allinstrumentation.enabled requires metrics.enabled. Without it, instrumentation stays off and a warning is logged, since nothing would expose the metrics.
3. Apply Instrumentation
Choose one of two approaches:
Option A: Instrument everything
from mitsuki import Application
from mitsuki.core.instrumentation import Instrumented
@Instrumented()
@Application
class App:
passOption B: Instrument specific components
from mitsuki import Service
from mitsuki.core.instrumentation import Instrumented
@Instrumented()
@Service()
class UserService:
async def get_user(self, user_id: int):
return user4. View Metrics
Start your application and access:
- Human-readable:
http://localhost:8000/metrics - Prometheus format:
http://localhost:8000/metrics/prometheus
Instrumentation Approaches
Application-Level Instrumentation
Apply @Instrumented() to your @Application class to automatically instrument all components:
from mitsuki import Application
from mitsuki.core.instrumentation import Instrumented
@Instrumented()
@Application
class App:
passThis automatically wraps all public methods in:
@Serviceclasses@Repositoryclasses@RestControllerclasses
What gets instrumented:
# All of these are automatically instrumented:
@Service()
class UserService:
async def get_user(self, user_id: int): # ✓ Tracked
return user
def _internal_helper(self): # ✗ Skipped (private method)
pass
@Repository()
class UserRepository:
async def find_by_email(self, email: str): # ✓ Tracked
return user
@RestController("/api/users")
class UserController:
@GetMapping("/{user_id}")
async def get_user(self, user_id: int): # ✓ Tracked (HTTP + component)
return userComponent-Level Instrumentation
Apply @Instrumented() to individual components for fine-grained control:
from mitsuki import Service, Repository
from mitsuki.core.instrumentation import Instrumented
# Only this service is instrumented
@Instrumented()
@Service()
class OrderService:
async def create_order(self, data: dict):
return order
# This service is NOT instrumented
@Service()
class EmailService:
async def send_email(self, to: str, subject: str):
passSelective instrumentation:
from mitsuki import Service
from mitsuki.core.instrumentation import Instrumented
# Instrument critical business logic
@Instrumented()
@Service()
class PaymentService:
async def process_payment(self, amount: float): # Tracked
pass
# Don't instrument high-frequency background tasks
@Service()
class CacheWarmer:
async def warm_cache(self): # Not tracked
passDisabling Instrumentation
Explicitly disable instrumentation on a component:
@Instrumented(enabled=False)
@Service()
class NoInstrumentationService:
async def fast_operation(self): # Not tracked
passOr disable globally in configuration:
instrumentation:
enabled: falseWhat Gets Tracked
HTTP Metrics (automatic for all HTTP requests):
- Request count by method, route and status code
- Response time distribution (histograms for percentiles)
The path label holds the matched route template (/users/{user_id}), not the requested URL, so metric cardinality stays bounded by the size of the route table. Requests that match no route are labelled <unmatched>.
Component Metrics (for instrumented components):
- Method call count (success vs failure), labelled by component and method
- Execution time per method
- Error rates
Scheduler Metrics (automatic when scheduler is enabled):
- Task execution count (success vs failure)
- Task execution duration
- Number of running tasks
System Metrics (automatic when instrumentation is enabled):
- CPU usage percentage
- Memory usage (RSS, VMS)
- Python traced memory, current and peak — only when
track_memory: true
Note: Private methods (starting with _) are never instrumented.
Which classes can be instrumented: @Service, @Repository, @CrudRepository and @RestController classes. @Instrumented has no effect on a plain @Component.
For a @CrudRepository, every method is recorded under the repository's name: the built-in methods (save, find_by_id, find_all, ...), query methods implemented from their names (find_by_email), @Query methods and methods you implement yourself:
@CrudRepository(entity=User)
class UserRepository:
async def find_by_email(self, email: str) -> Optional[User]: ... # ✓ Tracked
async def count_active(self) -> int: # ✓ Tracked
return len(await self.find_by_active(True)) # ✓ Tracked as wellInheritance: a call is recorded once, under the class of the instance it was made on. If an instrumented component inherits from another instrumented component, calls made through the subclass are labelled with the subclass name. A subclass that opts out with @Instrumented(enabled=False) is still recorded for the methods it inherits from an instrumented parent, since those calls go through the parent's instrumentation.
Metrics endpoints: requests to /metrics and /metrics/prometheus are not recorded, so scrapes never appear in HTTP or component metrics.
Metrics Endpoints
/metrics (Mitsuki Format)
Human-readable JSON with computed aggregations:
curl http://localhost:8000/metrics | jqExample response:
{
"enabled": true,
"timestamp": "2025-12-15T10:30:00.000000",
"instrumentation": {
"system": {
"memory": {"rss_mb": 85.2, "vms_mb": 150.3},
"cpu": {"percent": 1.2}
},
"http": {
"total_requests": 150,
"requests_by_method": {"GET": 100, "POST": 50},
"responses_by_status": {"200": 140, "404": 10},
"latency": {"avg_ms": 8.5}
},
"components": {
"UserService": {
"calls": 120,
"avg_duration_ms": 5.2,
"methods": {
"get_user": {"calls": 100, "avg_duration_ms": 4.1},
"create_user": {"calls": 20, "avg_duration_ms": 10.7}
}
}
}
}
}/metrics/prometheus (Prometheus Format)
Flat text format compatible with Prometheus scraping:
curl http://localhost:8000/metrics/prometheusExample response:
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",path="/api/users/{user_id}",status="200"} 100.0
# HELP http_request_duration_seconds HTTP request duration
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{method="GET",path="/api/users/{user_id}",le="0.005"} 50
http_request_duration_seconds_bucket{method="GET",path="/api/users/{user_id}",le="0.01"} 80
http_request_duration_seconds_sum{method="GET",path="/api/users/{user_id}"} 0.75
http_request_duration_seconds_count{method="GET",path="/api/users/{user_id}"} 100
# HELP component_calls_total Component method calls
# TYPE component_calls_total counter
component_calls_total{component="UserService",method="get_user",status="success"} 100.0
# HELP scheduler_task_executions_total Total number of scheduled task executions
# TYPE scheduler_task_executions_total counter
scheduler_task_executions_total{task="BackgroundService.cleanup",status="success"} 42.0
# HELP scheduler_task_duration_seconds Scheduled task execution duration in seconds
# TYPE scheduler_task_duration_seconds histogram
scheduler_task_duration_seconds_bucket{task="BackgroundService.cleanup",le="0.005"} 10
scheduler_task_duration_seconds_sum{task="BackgroundService.cleanup"} 5.25
scheduler_task_duration_seconds_count{task="BackgroundService.cleanup"} 42Scheduler Metrics
When the scheduler and metrics are enabled, Mitsuki automatically records metrics for all @Scheduled tasks. Scheduler metrics need only scheduler.enabled and metrics.enabled; instrumentation.enabled governs HTTP and component instrumentation and is not required here.
Configuration:
scheduler:
enabled: true
metrics:
enabled: trueMetrics tracked:
scheduler_task_executions_total- Counter with labels{task, status}scheduler_task_duration_seconds- Histogram with label{task}scheduler_tasks_running- Gauge with label{task}, counting executions currently in flight
Example queries:
# Task execution rate
rate(scheduler_task_executions_total{task="BackgroundService.cleanup"}[5m])
# Task failure rate
rate(scheduler_task_executions_total{status="failure"}[5m]) / rate(scheduler_task_executions_total[5m])
# Average task duration
rate(scheduler_task_duration_seconds_sum[5m]) / rate(scheduler_task_duration_seconds_count[5m])
# P95 task duration
histogram_quantile(0.95, sum(rate(scheduler_task_duration_seconds_bucket[5m])) by (le, task))
# Currently running tasks
scheduler_tasks_runningThese metrics are automatically included in both /metrics and /metrics/prometheus endpoints.
For more information on creating scheduled tasks, see Scheduled Tasks.
Custom Metrics
Track application-specific operational data using InstrumentationProvider.
Basic Usage
Inject InstrumentationProvider into your component:
from mitsuki import Service
from mitsuki.core.instrumentation import InstrumentationProvider
@Service()
class OrderService:
def __init__(self, instrumentation: InstrumentationProvider):
self.instrumentation = instrumentation
async def create_order(self, user_id: int):
order = await self.save_order(user_id)
# Record custom metric
self.instrumentation.record_metric(
metric_name="orders_created_total",
value=1,
labels={"source": "api"}
)
return orderRecording Metrics
The record_metric method accepts:
metric_name(str): Metric identifiervalue(float): Value to recordlabels(dict): Dimensional data for filtering/grouping
All custom metrics are stored as counters (monotonically increasing values).
Custom metrics are exposed only at /metrics/prometheus. The JSON summary at /metrics covers the built-in metrics and does not include them.
Metric names must match [a-zA-Z_:][a-zA-Z0-9_:]* and label names [a-zA-Z_][a-zA-Z0-9_]* (label names starting with __ are reserved). A name already used by a built-in gauge or histogram cannot be reused for a counter. record_metric raises ValueError for any of these, even while metrics are disabled, since a single invalid name would make Prometheus reject the entire scrape.
NOTE
The record_metric method is designed for counters only (e.g., counting events). It is not suitable for tracking durations or other values where you would need averages or percentiles. For those use cases, direct integration with a metrics library would be required.
Examples
Track database operations:
@Service()
class ProductService:
def __init__(self, repo: ProductRepository, instrumentation: InstrumentationProvider):
self.repo = repo
self.instrumentation = instrumentation
async def create_product(self, data: dict):
product = await self.repo.save(data)
# Track write operations
self.instrumentation.record_metric(
metric_name="database_writes_total",
value=1,
labels={"table": "products", "operation": "insert"}
)
return product
async def get_all_products(self):
products = await self.repo.find_all()
# Track query result size
self.instrumentation.record_metric(
metric_name="database_query_rows_returned",
value=len(products),
labels={"table": "products", "query_type": "find_all"}
)
# Flag expensive full table scans
self.instrumentation.record_metric(
metric_name="database_full_scan_total",
value=1,
labels={"table": "products"}
)
return productsTrack cache performance:
@Service()
class CacheService:
def __init__(self, instrumentation: InstrumentationProvider):
self.instrumentation = instrumentation
self.cache = {}
async def get(self, key: str):
if key in self.cache:
# Cache hit
self.instrumentation.record_metric(
metric_name="cache_operations_total",
value=1,
labels={"operation": "hit", "cache_name": "user_cache"}
)
return self.cache[key]
# Cache miss
self.instrumentation.record_metric(
metric_name="cache_operations_total",
value=1,
labels={"operation": "miss", "cache_name": "user_cache"}
)
return NoneTrack external API calls:
@Service()
class PaymentService:
def __init__(self, instrumentation: InstrumentationProvider):
self.instrumentation = instrumentation
async def charge_card(self, amount: float):
try:
response = await self.stripe_api.charge(amount)
# Track successful API call
self.instrumentation.record_metric(
metric_name="external_api_calls_total",
value=1,
labels={"service": "stripe", "status": "success"}
)
return response
except Exception as e:
# Track failed API call
self.instrumentation.record_metric(
metric_name="external_api_calls_total",
value=1,
labels={"service": "stripe", "status": "failure"}
)
raiseTrack business events:
@Service()
class SubscriptionService:
def __init__(self, instrumentation: InstrumentationProvider):
self.instrumentation = instrumentation
async def upgrade_subscription(self, user_id: int, plan: str):
# Perform upgrade logic...
# Track subscription changes
self.instrumentation.record_metric(
metric_name="subscription_changes_total",
value=1,
labels={"action": "upgrade", "plan": plan}
)Configuration
Basic Configuration
instrumentation:
enabled: true # Enable/disable instrumentation
track_memory: false # Also collect Python traced memory (tracemalloc)
metrics:
enabled: true # Enable metrics endpoints
path: /metrics # Base path for metrics
allowed_ips: [] # Empty = allow all IPsIP Allowlisting
Restrict access to metrics endpoints by IP:
Reverse Proxies and Security
The check uses the client address the server reports (request.client.host). Behind a reverse proxy, load balancer or cloud gateway (like Nginx or an AWS ALB), Granian, the default server, reports the proxy as the client, so external requests look internal: allowlisting the proxy's address (e.g. 127.0.0.1) exposes the endpoints to everyone the proxy serves. Uvicorn fills the client from X-Forwarded-For instead.
In a proxied setup, don't allowlist the proxy. Restrict the metrics paths at the proxy, or scrape the application on an address the proxy doesn't route to.
metrics:
enabled: true
allowed_ips:
- "127.0.0.1" # Localhost
- "10.0.0.0/8" # Private network
- "172.16.0.0/12" # Docker networks (172.16-31.x.x)
- "192.168.0.0/16" # Home/office networkEntries are parsed when the application starts; an invalid entry fails startup with an error naming it. Addresses are compared as addresses, so ::ffff:127.0.0.1 (an IPv4 client on a dual-stack listener) matches 127.0.0.1. When allowed_ips is non-empty, a client whose address the server does not report, or reports as something other than an IP, is denied.
When access is denied:
- A warning is logged:
Metrics access denied for IP: 192.168.1.100 - The client gets exactly the response a path that doesn't exist gets: a
404with the bodyNot Found, for every HTTP method. The check runs before routing, so a denied client can't tell that the metrics endpoints exist.
Environment Variables
Override configuration via environment variables:
export INSTRUMENTATION_ENABLED=true
export INSTRUMENTATION_TRACK_MEMORY=false
export METRICS_ENABLED=true
export METRICS_PATH=/metricsMultiple Workers
Metrics are stored in the memory of the process that records them. With server.workers greater than 1, each worker keeps its own separate counters, histograms and gauges, and each request to /metrics or /metrics/prometheus is answered by whichever worker receives it. Consecutive Prometheus scrapes can therefore hit different workers, which Prometheus reads as counters jumping up and down or resetting, and rate() results become unreliable.
This is a current limitation. Until metrics are aggregated across workers, run instrumented applications with a single worker per process and scale by running more processes, each scraped as its own Prometheus target:
server:
workers: 1Scheduled tasks have the same per-worker behaviour; see Multi-Worker Considerations.
Memory Tracking
Whenever instrumentation is enabled, Mitsuki samples process CPU and memory (RSS and VMS) via psutil every 5 seconds. These are cheap and always collected.
track_memory additionally reports Python's traced memory (current and peak) via tracemalloc. This adds significant allocation overhead, so it defaults to false and should stay off unless you are actively investigating memory.
instrumentation:
track_memory: falseIntegration with Prometheus & Grafana
Prometheus Scraping
Configure Prometheus to scrape your Mitsuki application:
prometheus.yml:
scrape_configs:
- job_name: 'mitsuki'
scrape_interval: 15s
static_configs:
- targets: ['localhost:8000']
metrics_path: '/metrics/prometheus'Grafana Dashboard
Mitsuki ships a ready-made Grafana dashboard covering HTTP, component (including per-method breakdowns), scheduler and system metrics. Write it out with the CLI:
# Write dashboard.json into the current directory
mitsuki grafana-dashboard
# Or into a specific directory (created if needed)
mitsuki grafana-dashboard -o ./grafana/dashboards/Point a Grafana dashboard provider at the output directory, or import the file through the Grafana UI. The dashboard queries a Prometheus datasource scraping /metrics/prometheus.
For a fully wired setup that generates and provisions the dashboard automatically, see the instrumentation demo.
Grafana Queries
While it's up to you to write queries useful for your particular use-case, here are a few generic examples:
Request rate:
rate(http_requests_total[5m])Average response time:
rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m])P95 latency:
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))Error rate:
rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m])Component call rate:
rate(component_calls_total{component="UserService"}[5m])Custom metric (database writes):
rate(database_writes_total{table="users"}[5m])Scheduler task execution rate:
rate(scheduler_task_executions_total[5m])Scheduler task failure rate:
rate(scheduler_task_executions_total{status="failure"}[5m]) / rate(scheduler_task_executions_total[5m])Docker Compose Example
version: '3.8'
services:
app:
build: .
ports:
- "8000:8000"
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
environment:
# For illustrative purposes, anonymous mode
- GF_AUTH_ANONYMOUS_ENABLED=true
- GF_AUTH_ANONYMOUS_ORG_ROLE=AdminComplete Example
For a full working example, check out the examples/instrumentation_demo app:
cd examples/instrumentation_demo
docker compose up -d --buildIncludes:
- Application-level instrumentation setup
- Custom operational metrics
- Pre-configured Grafana dashboards
- Prometheus + Grafana integration
Next Steps
- Scheduled Tasks - Learn about creating scheduled tasks
- Configuration - Learn more about
application.yml - Decorators - Understand
@Service,@Repository,@RestController - Controllers - Build HTTP endpoints