The Distribution-Aware Encoder is a powerful preprocessing layer that automatically detects and handles various data distributions. It intelligently transforms your data while preserving its statistical properties, leading to better model performance.
from kdp import PreprocessingModel, FeatureType
# Define numerical features
features_specs = {
"price": FeatureType.FLOAT_NORMALIZED,
"volume": FeatureType.FLOAT_RESCALED,
"rating": FeatureType.FLOAT_NORMALIZED
}
# Initialize model with distribution-aware encoding
preprocessor = PreprocessingModel(
path_data="data/my_data.csv",
features_specs=features_specs,
use_distribution_aware=True, # Enable distribution-aware encoding
distribution_aware_bins=1000 # Number of bins for distribution analysis
)from kdp import FeatureType, PreprocessingModel
from kdp.features import NumericalFeature
from kdp.layers.distribution_aware_encoder_layer import DistributionType
features_specs = {
"price": NumericalFeature(
name="price",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.LOG_NORMAL # Specify distribution
),
"volume": NumericalFeature(
name="volume",
feature_type=FeatureType.FLOAT_RESCALED,
preferred_distribution=DistributionType.ZERO_INFLATED # Handle sparse data
)
}
preprocessor = PreprocessingModel(
path_data="data/my_data.csv",
features_specs=features_specs,
use_distribution_aware=True,
distribution_aware_bins=1000,
)| Distribution Type | Description | Detection Criteria | Use Case |
|---|---|---|---|
| Normal | Standard bell curve | Skewness < 0.5, Kurtosis ≈ 3.0 | Height, weight measurements |
| Heavy-Tailed | Longer tails than normal | Kurtosis > 4.0 | Financial returns |
| Multimodal | Multiple peaks | Multiple histogram peaks | Mixed populations |
| Uniform | Even distribution | Bounded between 0 and 1 | Random sampling |
| Exponential | Exponential decay | Positive values, skewness > 1.0 | Time between events |
| Log-Normal | Normal after log transform | Positive values, skewness > 2.0 | Income distribution |
| Discrete | Finite distinct values | Low unique value ratio (< 0.1) | Count data |
| Periodic | Cyclic patterns | Significant autocorrelation | Seasonal data |
| Sparse | Many zeros | Zero ratio > 0.5 | User activity data |
| Beta | Bounded with shape parameters | Bounded [0,1], skewness > 0.5 | Proportions |
| Gamma | Positive, right-skewed | Positive values, mild skewness | Waiting times |
| Poisson | Count data | Discrete positive values | Event counts |
| Cauchy | Extremely heavy-tailed | Very high kurtosis (> 10.0) | Extreme events |
| Zero-Inflated | Excess zeros | Moderate zero ratio (0.3-0.5) | Rare events |
| Bounded | Known bounds | Explicit bounds provided | Physical measurements |
| Ordinal | Ordered categories | Discrete ordered values | Ratings, scores |
These two are the whole model-level surface:
| Parameter | Type | Default | Description |
|---|---|---|---|
use_distribution_aware |
bool | False | Enable distribution-aware encoding |
distribution_aware_bins |
int | 1000 | Number of bins for distribution analysis |
Per feature, set preferred_distribution on a NumericalFeature to skip
automatic detection and force a specific distribution.
!!! warning "Automatic detection is a heuristic, and it is not always right"
Measured over six seeded samples, detection identifies heavy_tailed,
log_normal, discrete, periodic, sparse and beta correctly and
consistently. Four shapes are consistently confused with a neighbour:
| Actual | Detected as |
|---|---|
| `normal` | `multimodal` |
| `uniform` | `multimodal` |
| `multimodal` | `periodic` |
| `exponential` | `log_normal` |
The encoding still works — a neighbouring distribution's transform is
usually a reasonable choice — but if a column's shape matters to you,
set `preferred_distribution` on the feature rather than relying on
detection. `test/layers/test_distribution_aware_encoder.py` pins this
behaviour, so it cannot change without being noticed.
transform_type accepts these; anything else raises. auto is the default on
the encoder and picks from the rest using the column's own shape —
whether it is bounded, contains zeros, or contains negative values.
| Name | Requires | Use for |
|---|---|---|
auto | — | Let KDP choose from the column's shape. |
none | — | Pass the values through unchanged. |
log | Strictly positive | Long right tails, such as income. |
sqrt | Non-negative | Milder right skew; tolerates zeros. |
cube-root | Any sign | Skew in data that also goes negative. |
arcsinh | Any sign | Log-like compression that accepts zero and negatives. |
box-cox | Strictly positive | Power transform toward normality. |
yeo-johnson | Any sign | Box-Cox for data including zero and negatives. |
logit | Values inside (0, 1) | Proportions and rates. |
min-max | — | Rescale to a fixed range. |
robust-scale | — | Median and IQR per feature; resists outliers. |
quantile | — | Rank-based, per feature; flattens any shape. |
!!! tip "auto narrows the candidates to what the data allows"
A column with zeros never gets log, and one with negative values never
gets box-cox, so auto cannot produce infinities from an unsuitable
transform. Restrict the search yourself with
auto_candidates=["log", "sqrt"].
The processor builds this layer for you with detect_periodicity=True,
handle_sparsity=True, adaptive_binning=True and mixture_components=3
fixed. To change them, use the layer directly through a
custom preprocessing pipeline — passing them
to PreprocessingModel raises TypeError.
| Parameter | Type | Default | Description |
|---|---|---|---|
detect_periodicity |
bool | True | Detect and handle periodic patterns |
handle_sparsity |
bool | True | Special handling for sparse data |
adaptive_binning |
bool | None | Learn bin boundaries from the batch rather than using fixed ones |
embedding_dim |
int | None | Output dimension for feature projection |
add_distribution_embedding |
bool | False | Add learned distribution type embedding |
epsilon |
float | 1e-6 | Small value to prevent numerical issues |
transform_type |
str | "auto" | Type of transformation to apply |
- Start with automatic detection
- Specify preferred distributions only when confident
- Use appropriate bin sizes for your data scale
- Monitor detection accuracy with known distributions
- Enable periodic detection for time series data
- Use sparse handling for data with many zeros
- Consider memory usage with large bin sizes
- Use appropriate embedding dimensions
from kdp import PreprocessingModel, FeatureType
from kdp.features import NumericalFeature
from kdp.layers.distribution_aware_encoder_layer import DistributionType
# Define financial features
features_specs = {
"price": NumericalFeature(
name="price",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.LOG_NORMAL
),
"volume": NumericalFeature(
name="volume",
feature_type=FeatureType.FLOAT_RESCALED,
preferred_distribution=DistributionType.ZERO_INFLATED
),
"volatility": NumericalFeature(
name="volatility",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.CAUCHY
)
}
# Create preprocessing model
preprocessor = PreprocessingModel(
path_data="data/financial_data.csv",
features_specs=features_specs,
use_distribution_aware=True,
distribution_aware_bins=1000,
embedding_dim=32, # Project to fixed dimension
)</div>
from kdp import DistributionType, FeatureType, NumericalFeature, PreprocessingModel
features_specs = {
"temperature": NumericalFeature(
name="temperature",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.NORMAL
),
"humidity": NumericalFeature(
name="humidity",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.BETA # Bounded between 0-100%
),
"pressure": NumericalFeature(
name="pressure",
feature_type=FeatureType.FLOAT_NORMALIZED,
preferred_distribution=DistributionType.NORMAL
)
}
preprocessor = PreprocessingModel(
path_data="data/sensor_data.csv",
features_specs=features_specs,
use_distribution_aware=True,
distribution_aware_bins=500, # Fewer bins for simpler distributions
embedding_dim=16 # Smaller embedding for simpler patterns
)</div>
The distribution-aware encoder architecture automatically adapts to your data's distribution, transforming numerical features to better match their underlying statistical properties, improving model performance.
<style> /* Base styling */ body { font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, Helvetica, Arial, sans-serif; line-height: 1.6; color: #333; margin: 0; padding: 0; } /* Feature header */ .feature-header { background: linear-gradient(135deg, #4caf50 0%, #8bc34a 100%); border-radius: 10px; padding: 30px; margin: 30px 0; box-shadow: 0 4px 6px rgba(0,0,0,0.1); color: white; } .feature-title h2 { margin-top: 0; font-size: 28px; } .feature-title p { font-size: 18px; margin-bottom: 0; opacity: 0.9; } /* Overview card */ .overview-card { background-color: #fff; border-radius: 10px; padding: 20px 25px; margin: 20px 0; box-shadow: 0 2px 5px rgba(0,0,0,0.05); border-left: 4px solid #4caf50; } .overview-card p { margin: 0; font-size: 16px; } /* Key benefits */ .key-benefits { display: grid; grid-template-columns: repeat(auto-fill, minmax(250px, 1fr)); gap: 20px; margin: 30px 0; } .benefit-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .benefit-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .benefit-icon { font-size: 2.5em; margin-bottom: 15px; } .benefit-card h3 { margin: 0 0 10px 0; color: #4caf50; } .benefit-card p { margin: 0; } /* Use cases */ .use-cases-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(200px, 1fr)); gap: 20px; margin: 30px 0; } .use-case-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; display: flex; flex-direction: column; align-items: center; text-align: center; } .use-case-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .use-case-icon { font-size: 2.5em; margin-bottom: 15px; } .use-case-card h3 { margin: 0 0 10px 0; color: #4caf50; } .use-case-card p { margin: 0; } /* Code containers */ .code-container { background-color: #f8f9fa; border-radius: 8px; overflow: hidden; box-shadow: 0 2px 5px rgba(0,0,0,0.1); margin: 20px 0; } .code-container pre { margin: 0; padding: 20px; } /* Tables */ .table-container { margin: 30px 0; border-radius: 10px; overflow: hidden; box-shadow: 0 4px 8px rgba(0,0,0,0.05); } .distributions-table, .config-table { width: 100%; border-collapse: collapse; } .distributions-table th, .config-table th { background-color: #e8f5e9; padding: 15px; text-align: left; font-weight: 600; border-bottom: 2px solid #4caf50; } .distributions-table td, .config-table td { padding: 12px 15px; border-bottom: 1px solid #eaecef; } .distributions-table tr:nth-child(even), .config-table tr:nth-child(even) { background-color: #f8f9fa; } .distributions-table tr:hover, .config-table tr:hover { background-color: #e8f5e9; } /* Best practices */ .best-practices-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(300px, 1fr)); gap: 20px; margin: 30px 0; } .best-practice-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; } .best-practice-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .best-practice-card h3 { margin-top: 0; color: #4caf50; } .best-practice-card ul { margin: 0; padding-left: 20px; } .best-practice-card li { margin-bottom: 5px; } /* Examples */ .examples-container { display: grid; grid-template-columns: repeat(auto-fill, minmax(400px, 1fr)); gap: 20px; margin: 30px 0; } .example-card { background-color: #fff; border-radius: 10px; padding: 20px; box-shadow: 0 4px 8px rgba(0,0,0,0.05); transition: transform 0.3s ease, box-shadow 0.3s ease; } .example-card:hover { transform: translateY(-5px); box-shadow: 0 8px 16px rgba(0,0,0,0.1); } .example-card h3 { margin-top: 0; color: #4caf50; } /* Architecture diagram */ .architecture-diagram { background-color: white; border-radius: 10px; padding: 20px; margin: 30px 0; box-shadow: 0 4px 8px rgba(0,0,0,0.05); text-align: center; } .architecture-image { max-width: 100%; border-radius: 5px; } .diagram-caption { margin-top: 20px; text-align: center; font-style: italic; } /* Related topics */ .related-topics { display: flex; flex-wrap: wrap; gap: 15px; margin: 30px 0; } .topic-link { display: flex; align-items: center; padding: 10px 15px; background-color: #e8f5e9; border-radius: 8px; text-decoration: none; color: #333; box-shadow: 0 2px 5px rgba(0,0,0,0.05); transition: background-color 0.3s ease, transform 0.3s ease; } .topic-link:hover { background-color: #c8e6c9; transform: translateY(-2px); } .topic-icon { font-size: 1.2em; margin-right: 10px; } /* Navigation */ .nav-container { display: flex; justify-content: space-between; margin: 40px 0; } .nav-button { display: flex; align-items: center; padding: 10px 15px; background-color: #f8f9fa; border-radius: 8px; text-decoration: none; color: #333; box-shadow: 0 2px 5px rgba(0,0,0,0.1); transition: background-color 0.3s ease, transform 0.3s ease; } .nav-button:hover { background-color: #e8f5e9; transform: translateY(-2px); } .nav-button.prev { padding-left: 10px; } .nav-button.next { padding-right: 10px; } .nav-icon { font-size: 1.2em; margin: 0 8px; } /* Responsive adjustments */ @media (max-width: 768px) { .key-benefits, .use-cases-container, .best-practices-container, .examples-container { grid-template-columns: 1fr; } .related-topics { flex-direction: column; } } </style>