cano-collector

command module
v0.0.0-...-82631bb Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Nov 30, 2025 License: Apache-2.0 Imports: 24 Imported by: 0

README ยถ

cano-collector

SonarQube Cloud

Quality Gate Status Bugs Code Smells Coverage

Cano Collector

Cano Collector is an open-source alert and event ingestion agent for Kubernetes, designed to help developers and DevOps teams better understand incidents in their clusters by enriching raw alerts and events with valuable context.

Cano Collector is part of the broader Kubecano platform. It runs on Kubernetes clusters and connects telemetry data with notifications, enrichment pipelines, and (in future releases) AI-based analysis.


๐Ÿš€ Why Cano Collector?

Traditional alerts and crash loops often lack the full story. Cano Collector gives you:

  • Deep context behind alerts and events
  • Flexible routing to the right teams
  • Rich formatting with structured data and attachments
  • Unified view of Kubernetes health, enriched and sent where it matters

Whether it's an OOMKilled pod or a CrashLoopBackOff, Cano Collector helps you understand why something broke โ€” not just that it did.


๐Ÿงฉ What It Does

Cano Collector listens for Kubernetes cluster signals, including:

  • ๐Ÿ“ฃ Alerts from Alertmanager
  • โš ๏ธ Kubernetes Events such as:
    • Pod restarts / CrashLoops
    • Helm release failures
    • Resource quota violations

For each alert or event, Cano Collector:

  1. Builds a structured Issue object that includes:

    • Type: alert or event
    • Source: prometheus, k8s, helm, etc.
    • Severity: HIGH, LOW, INFO, DEBUG
    • Timestamps: created/started/resolved
  2. Enriches it with context through Enrichment blocks:

    • Pod logs as MarkdownBlock
    • Resource configuration as TableBlock
    • File attachments as FileBlock
    • Structured data as JsonBlock
  3. Sends enriched data to configured destinations:

    • ๐Ÿ’ฌ Slack channels (MVP - Available Now)
    • ๐Ÿงญ Kubecano SaaS (Planned)
    • ๐Ÿ“Ÿ PagerDuty, OpsGenie (Planned)
    • ๐Ÿ”€ Kafka topics (Planned)

๐Ÿ“ฆ Architecture Overview

Cano Collector follows a clean architecture pattern with clear separation of concerns:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   Alertmanager  โ”‚    โ”‚   Kubernetes    โ”‚    โ”‚   Other Sources โ”‚
โ”‚   (Prometheus)  โ”‚    โ”‚     Events      โ”‚    โ”‚                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
          โ”‚                      โ”‚                      โ”‚
          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚     Cano Collector        โ”‚
                    โ”‚  (Deployed on K8s)        โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                  โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚      Destination          โ”‚
                    โ”‚   (Strategy Pattern)      โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                  โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚       Sender              โ”‚
                    โ”‚   (Factory Pattern)       โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                  โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚    External Services      โ”‚
                    โ”‚   Slack, Teams, etc.      โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
Core Components
  • Issue: The central data structure containing alert/event information
  • Enrichment: Additional context blocks (logs, tables, files, etc.)
  • Destination: Strategy pattern implementation for different notification channels
  • Sender: Factory pattern implementation for API communication

๐ŸŽฏ Current Status (MVP)

โœ… Available Now (v0.0.24)
  • Alertmanager Integration: Full webhook-based alert processing from Prometheus/Alertmanager
  • Slack Integration: Full-featured Slack destination with:
    • Rich message formatting with blocks and attachments
    • Thread support for related alerts with intelligent grouping
    • File uploads for logs and structured data
    • Color-coded messages based on severity
    • Table formatting for structured data
    • Message deduplication and fingerprinting
  • Workflow System: Extensible alert processing with:
    • Built-in workflow actions (pod logs, issue enrichment)
    • Configurable triggers and conditions
    • Support for custom TypeScript workflows (planned)
  • Alert Enrichment: Automatic context gathering from Kubernetes:
    • Pod logs (current and previous containers)
    • Resource metadata and labels
    • Alert annotations and severity mapping
  • Observability: Production-ready monitoring with:
    • Health checks (liveness and readiness)
    • Prometheus metrics export
    • OpenTelemetry tracing support
    • Structured logging
๐Ÿ”ฎ Next Releases

Additional integrations and features will be added in future versions:

  • MS Teams Integration: Adaptive Cards support
  • Direct Kubernetes Event Processing: Watch and process K8s events (BackOff, CrashLoopBackOff, etc.)
  • Team-Based Routing: Multi-team alert distribution
  • PagerDuty Integration: Incident lifecycle management
  • Jira Integration: Ticket creation and management
  • Additional Senders: DataDog, Kafka, ServiceNow, OpsGenie

๐Ÿ”ง Installation

Prerequisites
  • Kubernetes cluster (1.19+)
  • Helm 3.x
  • Alertmanager configured (optional)
Quick Start
  1. Add the Helm repository:

    helm repo add kubecano https://kubecano.github.io/helm-charts
    helm repo update
    
  2. Create a values file (values.yaml):

    destinations:
      - name: slack_production
        type: slack
        params:
          webhook_url: "https://hooks.slack.com/services/YOUR/WEBHOOK/URL"
          channel: "#alerts"
          username: "Cano Collector"
    
  3. Install Cano Collector:

    helm install cano-collector kubecano/cano-collector -f values.yaml
    
Configuration

See the Configuration Guide for detailed setup instructions.


๐Ÿ“Œ Use Cases

Current Implementation (v0.0.24)

Cano Collector currently processes Prometheus alerts from Alertmanager and enriches them with Kubernetes context:

  • Pod CrashLoopBackOff Alerts:

    • Receive KubePodCrashLooping alerts in Slack
    • Automatic pod logs attachment (previous and current container)
    • Container restart count and exit codes
    • Resource configuration and limits
    • Color-coded severity and threaded conversations
  • General Kubernetes Alerts:

    • Any Prometheus/Alertmanager alert routed to cano-collector
    • Rich Slack formatting with alert metadata
    • Alert annotations and labels as enrichments
    • Resolved alerts posted as thread replies
  • Custom Alert Enrichment:

    • Configurable workflows for different alert types
    • Pod logs collection for container failures
    • Kubernetes resource metadata extraction
    • Extensible action system for custom enrichments
Planned Features (Future Releases)
  • Direct Kubernetes Event Processing: Watch and process K8s events in real-time (BackOff, ImagePull, Eviction, etc.)
  • Multi-Channel Routing: Send different alerts to different teams and destinations
  • Additional Destinations: MS Teams, PagerDuty, Jira, DataDog, Kafka, ServiceNow
  • Advanced Enrichments:
    • OOMKilled analysis with memory graphs
    • Resource usage trends
    • Node health correlations
    • Deployment history and changes
  • Team-Based Alert Distribution: Route alerts based on namespace, labels, and severity to specific teams
  • Incident Management Integration: Create tickets in Jira, incidents in PagerDuty with full lifecycle tracking

๐Ÿ—๏ธ Development

Architecture Documentation
Contributing

We welcome contributions! Please see our Contributing Guide for details.

Building
# Build the binary
go build -o cano-collector ./main.go

# Build the Docker image
docker build -t cano-collector .

# Run tests
go test ./...

๐Ÿ”ฎ Roadmap

โœ… Phase 1: MVP (Completed - v0.0.24)

Core alert processing and Slack integration:

  • โœ… Alertmanager webhook integration
  • โœ… Slack destination with full feature set
  • โœ… Workflow system with configurable actions
  • โœ… Basic alert enrichment (pod logs, metadata)
  • โœ… Health checks and metrics
  • โœ… OpenTelemetry tracing
๐Ÿšง Phase 2: Core Platform Features (In Progress)

Expanding destination support and processing capabilities:

  • ๐Ÿ”จ Direct Kubernetes Event Processing (watch K8s API events)
  • ๐Ÿ”จ MS Teams destination
  • ๐Ÿ”จ PagerDuty integration
  • ๐Ÿ”จ Enhanced workflow system
  • ๐Ÿ”จ Team-based routing
๐Ÿ“‹ Phase 3: Enterprise Features (Planned)

Advanced integrations and analysis:

  • ๐Ÿ“… Jira Service Management integration
  • ๐Ÿ“… OpsGenie integration
  • ๐Ÿ“… DataDog event correlation
  • ๐Ÿ“… Kafka streaming
  • ๐Ÿ“… Alert deduplication system
  • ๐Ÿ“… Async processing queue
๐ŸŒŸ Phase 4: Advanced Capabilities (Future)

Platform ecosystem and intelligence:

  • ๐ŸŒŸ Kubecano CLI tool
  • ๐ŸŒŸ Official Slack App
  • ๐ŸŒŸ Advanced monitoring and observability
  • ๐ŸŒŸ Custom TypeScript workflows (runtime)
  • ๐ŸŒŸ ServiceNow integration
๐Ÿš€ Phase 5: SaaS Platform (Long-term Vision)

Multi-tenant SaaS offering:

  • ๐Ÿš€ Web dashboard and SaaS platform
  • ๐Ÿš€ Multi-cluster management
  • ๐Ÿš€ AI-powered root cause analysis
  • ๐Ÿš€ Automated remediation suggestions
  • ๐Ÿš€ Advanced correlation and anomaly detection

Note: This roadmap is subject to change based on community feedback and priorities. Specific release dates are not provided as development is driven by community needs and contributions.


๐Ÿ‘ฅ Who is this for?

If you're a:

  • DevOps engineer managing production Kubernetes
  • Developer tired of vague alerts
  • SRE building observability tooling
  • Platform team looking for better incident response

โ€ฆthen Cano Collector is for you.


๐Ÿ“ฌ Get Involved

Join us in making Kubernetes incidents understandable!


๐Ÿ“ License

Cano Collector is licensed under the Apache 2.0 License.

Documentation ยถ

The Go Gopher

There is no documentation for this package.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL