Skip to content

Latest commit

 

History

History
902 lines (793 loc) · 23.2 KB

File metadata and controls

902 lines (793 loc) · 23.2 KB

🔄 Distribution-Aware Encoding

Distribution-Aware Encoding

Automatically detect and handle various data distributions for optimal preprocessing.

📋 Overview

The Distribution-Aware Encoder is a powerful preprocessing layer that automatically detects and handles various data distributions. It intelligently transforms your data while preserving its statistical properties, leading to better model performance.

🔍

Automatic Distribution Detection

Identifies data patterns using statistical analysis

⚙️

Smart Transformations

Applies distribution-specific preprocessing

🚀

Production-Ready

Built with pure TensorFlow operations for deployment

🔌

Flexible Integration

Works seamlessly with KDP's preprocessing pipeline

📊

Graph Mode Compatible

Works in both eager and graph execution modes

📦

Memory Efficient

Optimized for large-scale datasets

🎯 Use Cases

💰

Financial Data

Handling heavy-tailed distributions in price movements

📡

Sensor Data

Processing periodic patterns in time series

👤

User Behavior

Managing sparse data with many zeros

🌍

Natural Phenomena

Handling multimodal distributions

🔢

Count Data

Processing discrete and zero-inflated distributions

🚀 Getting Started

Basic Usage

from kdp import PreprocessingModel, FeatureType

# Define numerical features
features_specs = {
    "price": FeatureType.FLOAT_NORMALIZED,
    "volume": FeatureType.FLOAT_RESCALED,
    "rating": FeatureType.FLOAT_NORMALIZED
}

# Initialize model with distribution-aware encoding
preprocessor = PreprocessingModel(
    path_data="data/my_data.csv",
    features_specs=features_specs,
    use_distribution_aware=True,  # Enable distribution-aware encoding
    distribution_aware_bins=1000  # Number of bins for distribution analysis
)

Advanced Configuration

from kdp import FeatureType, PreprocessingModel

from kdp.features import NumericalFeature
from kdp.layers.distribution_aware_encoder_layer import DistributionType

features_specs = {
    "price": NumericalFeature(
        name="price",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.LOG_NORMAL  # Specify distribution
    ),
    "volume": NumericalFeature(
        name="volume",
        feature_type=FeatureType.FLOAT_RESCALED,
        preferred_distribution=DistributionType.ZERO_INFLATED  # Handle sparse data
    )
}

preprocessor = PreprocessingModel(
    path_data="data/my_data.csv",
    features_specs=features_specs,
    use_distribution_aware=True,
    distribution_aware_bins=1000,
)

📊 Supported Distributions

Distribution Type Description Detection Criteria Use Case
Normal Standard bell curve Skewness < 0.5, Kurtosis ≈ 3.0 Height, weight measurements
Heavy-Tailed Longer tails than normal Kurtosis > 4.0 Financial returns
Multimodal Multiple peaks Multiple histogram peaks Mixed populations
Uniform Even distribution Bounded between 0 and 1 Random sampling
Exponential Exponential decay Positive values, skewness > 1.0 Time between events
Log-Normal Normal after log transform Positive values, skewness > 2.0 Income distribution
Discrete Finite distinct values Low unique value ratio (< 0.1) Count data
Periodic Cyclic patterns Significant autocorrelation Seasonal data
Sparse Many zeros Zero ratio > 0.5 User activity data
Beta Bounded with shape parameters Bounded [0,1], skewness > 0.5 Proportions
Gamma Positive, right-skewed Positive values, mild skewness Waiting times
Poisson Count data Discrete positive values Event counts
Cauchy Extremely heavy-tailed Very high kurtosis (> 10.0) Extreme events
Zero-Inflated Excess zeros Moderate zero ratio (0.3-0.5) Rare events
Bounded Known bounds Explicit bounds provided Physical measurements
Ordinal Ordered categories Discrete ordered values Ratings, scores

⚙️ Configuration Options

On PreprocessingModel

These two are the whole model-level surface:

Parameter Type Default Description
use_distribution_aware bool False Enable distribution-aware encoding
distribution_aware_bins int 1000 Number of bins for distribution analysis

Per feature, set preferred_distribution on a NumericalFeature to skip automatic detection and force a specific distribution.

!!! warning "Automatic detection is a heuristic, and it is not always right" Measured over six seeded samples, detection identifies heavy_tailed, log_normal, discrete, periodic, sparse and beta correctly and consistently. Four shapes are consistently confused with a neighbour:

| Actual | Detected as |
|---|---|
| `normal` | `multimodal` |
| `uniform` | `multimodal` |
| `multimodal` | `periodic` |
| `exponential` | `log_normal` |

The encoding still works &mdash; a neighbouring distribution's transform is
usually a reasonable choice &mdash; but if a column's shape matters to you,
set `preferred_distribution` on the feature rather than relying on
detection. `test/layers/test_distribution_aware_encoder.py` pins this
behaviour, so it cannot change without being noticed.

Available transformations

transform_type accepts these; anything else raises. auto is the default on the encoder and picks from the rest using the column's own shape — whether it is bounded, contains zeros, or contains negative values.

Name Requires Use for
auto—Let KDP choose from the column's shape.
none—Pass the values through unchanged.
logStrictly positiveLong right tails, such as income.
sqrtNon-negativeMilder right skew; tolerates zeros.
cube-rootAny signSkew in data that also goes negative.
arcsinhAny signLog-like compression that accepts zero and negatives.
box-coxStrictly positivePower transform toward normality.
yeo-johnsonAny signBox-Cox for data including zero and negatives.
logitValues inside (0, 1)Proportions and rates.
min-max—Rescale to a fixed range.
robust-scale—Median and IQR per feature; resists outliers.
quantile—Rank-based, per feature; flattens any shape.

!!! tip "auto narrows the candidates to what the data allows" A column with zeros never gets log, and one with negative values never gets box-cox, so auto cannot produce infinities from an unsuitable transform. Restrict the search yourself with auto_candidates=["log", "sqrt"].

On the DistributionAwareEncoder layer

The processor builds this layer for you with detect_periodicity=True, handle_sparsity=True, adaptive_binning=True and mixture_components=3 fixed. To change them, use the layer directly through a custom preprocessing pipeline — passing them to PreprocessingModel raises TypeError.

Parameter Type Default Description
detect_periodicity bool True Detect and handle periodic patterns
handle_sparsity bool True Special handling for sparse data
adaptive_binning bool None Learn bin boundaries from the batch rather than using fixed ones
embedding_dim int None Output dimension for feature projection
add_distribution_embedding bool False Add learned distribution type embedding
epsilon float 1e-6 Small value to prevent numerical issues
transform_type str "auto" Type of transformation to apply

🎯 Best Practices

Distribution Detection

  • Start with automatic detection
  • Specify preferred distributions only when confident
  • Use appropriate bin sizes for your data scale
  • Monitor detection accuracy with known distributions

Performance Optimization

  • Enable periodic detection for time series data
  • Use sparse handling for data with many zeros
  • Consider memory usage with large bin sizes
  • Use appropriate embedding dimensions

Integration Tips

  • Combine with other KDP features for best results
  • Use appropriate feature types (FLOAT_NORMALIZED, FLOAT_RESCALED)
  • Monitor model performance with different configurations
  • Consider using distribution embeddings for complex patterns

🔍 Examples

Financial Data Processing

from kdp import PreprocessingModel, FeatureType
from kdp.features import NumericalFeature
from kdp.layers.distribution_aware_encoder_layer import DistributionType

# Define financial features
features_specs = {
    "price": NumericalFeature(
        name="price",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.LOG_NORMAL
    ),
    "volume": NumericalFeature(
        name="volume",
        feature_type=FeatureType.FLOAT_RESCALED,
        preferred_distribution=DistributionType.ZERO_INFLATED
    ),
    "volatility": NumericalFeature(
        name="volatility",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.CAUCHY
    )
}

# Create preprocessing model
preprocessor = PreprocessingModel(
    path_data="data/financial_data.csv",
    features_specs=features_specs,
    use_distribution_aware=True,
    distribution_aware_bins=1000,
    embedding_dim=32,        # Project to fixed dimension
)
</div>

Sensor Data Processing

from kdp import DistributionType, FeatureType, NumericalFeature, PreprocessingModel

features_specs = {
    "temperature": NumericalFeature(
        name="temperature",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.NORMAL
    ),
    "humidity": NumericalFeature(
        name="humidity",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.BETA  # Bounded between 0-100%
    ),
    "pressure": NumericalFeature(
        name="pressure",
        feature_type=FeatureType.FLOAT_NORMALIZED,
        preferred_distribution=DistributionType.NORMAL
    )
}

preprocessor = PreprocessingModel(
    path_data="data/sensor_data.csv",
    features_specs=features_specs,
    use_distribution_aware=True,
    distribution_aware_bins=500,  # Fewer bins for simpler distributions
    embedding_dim=16             # Smaller embedding for simpler patterns
)
</div>

📊 Model Architecture

Distribution-Aware Architecture

The distribution-aware encoder architecture automatically adapts to your data's distribution, transforming numerical features to better match their underlying statistical properties, improving model performance.

🔗 Related Topics


<style> /* Base styling */ body { font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, Helvetica, Arial, sans-serif; line-height: 1.6; color: #333; margin: 0; padding: 0; } /* Feature header */ .feature-header { background: linear-gradient(135deg, #4caf50 0%, #8bc34a 100%); border-radius: 10px; padding: 30px; margin: 30px 0; box-shadow: 0 4px 6px rgba(0,0,0,0.1); color: white; } .feature-title h2 { margin-top: 0; font-size: 28px; } .feature-title p { font-size: 18px; margin-bottom: 0; opacity: 0.9; } /* Overview card */ .overview-card { background-color: #fff; border-radius: 10px; padding: 20px 25px; margin: 20px 0; box-shadow: 0 2px 5px rgba(0,0,0,0.05); border-left: 4px solid #4caf50; } .overview-card p { margin: 0; font-size: 16px; } /* Key benefits */ .key-benefits { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .benefit-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .benefit-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .benefit-icon { font-size: 2.5em; margin-bottom: 15px; } .benefit-card h3 { margin: 0 0 10px 0; color: #4caf50; } .benefit-card p { margin: 0; } /* Use cases */ .use-cases-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(200px, 1fr)); gap: 20px; margin: 30px 0; } .use-case-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .use-case-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .use-case-icon { font-size: 2.5em; margin-bottom: 15px; } .use-case-card h3 { margin: 0 0 10px 0; color: #4caf50; } .use-case-card p { margin: 0; } /* Code containers */ .code-container { background-color: #f8f9fa; border-radius: 8px; overflow: hidden; box-shadow: 0 2px 5px rgba(0,0,0,0.1); margin: 20px 0; } .code-container pre { margin: 0; padding: 20px; } /* Tables */ .table-container { margin: 30px 0; border-radius: 10px; overflow: hidden; box-shadow: 0 4px 8px rgba(0,0,0,0.05); } .distributions-table, .config-table { width: 100%; border-collapse: collapse; } .distributions-table th, .config-table th { background-color: #e8f5e9; padding: 15px; text-align: left; font-weight: 600; border-bottom: 2px solid #4caf50; } .distributions-table td, .config-table td { padding: 12px 15px; border-bottom: 1px solid #eaecef; } .distributions-table tr:nth-child(even), .config-table tr:nth-child(even) { background-color: #f8f9fa; } .distributions-table tr:hover, .config-table tr:hover { background-color: #e8f5e9; } /* Best practices */ .best-practices-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(300px, 1fr)); gap: 20px; margin: 30px 0; } .best-practice-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; } .best-practice-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .best-practice-card h3 { margin-top: 0; color: #4caf50; } .best-practice-card ul { margin: 0; padding-left: 20px; } .best-practice-card li { margin-bottom: 5px; } /* Examples */ .examples-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(400px, 1fr)); gap: 20px; margin: 30px 0; } .example-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; } .example-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .example-card h3 { margin-top: 0; color: #4caf50; } /* Architecture diagram */ .architecture-diagram { background-color: white; border-radius: 10px; padding: 20px; margin: 30px 0; box-shadow: 0 4px 8px rgba(0,0,0,0.05); text-align: center; } .architecture-image { max-width: 100%; border-radius: 5px; } .diagram-caption { margin-top: 20px; text-align: center; font-style: italic; } /* Related topics */ .related-topics { display: flex; flex-wrap: wrap; gap: 15px; margin: 30px 0; } .topic-link { display: flex; align-items: center; padding: 10px 15px; background-color: #e8f5e9; border-radius: 8px; text-decoration: none; color: #333; box-shadow: 0 2px 5px rgba(0,0,0,0.05); transition: background-color 0.3s ease, transform 0.3s ease; } .topic-link:hover { background-color: #c8e6c9; transform: translateY(-2px); } .topic-icon { font-size: 1.2em; margin-right: 10px; } /* Navigation */ .nav-container { display: flex; justify-content: space-between; margin: 40px 0; } .nav-button { display: flex; align-items: center; padding: 10px 15px; background-color: #f8f9fa; border-radius: 8px; text-decoration: none; color: #333; box-shadow: 0 2px 5px rgba(0,0,0,0.1); transition: background-color 0.3s ease, transform 0.3s ease; } .nav-button:hover { background-color: #e8f5e9; transform: translateY(-2px); } .nav-button.prev { padding-left: 10px; } .nav-button.next { padding-right: 10px; } .nav-icon { font-size: 1.2em; margin: 0 8px; } /* Responsive adjustments */ @media (max-width: 768px) { .key-benefits, .use-cases-container, .best-practices-container, .examples-container { grid-template-columns: 1fr; } .related-topics { flex-direction: column; } } </style>