A comprehensive end-to-end pipeline for discovering, cleaning, and annotating industrial time series datasets for machine learning.
The fastest way to install a skill directly from the registry.
npx clawhub@latest install data-cleaning-annotation-workflow
Copy the skill folder to one of these locations
~/.openclaw/skills/ <project>/skills/ Priority: Workspace > Local > Bundled
Copy this prompt to OpenClaw to install it automatically.
Help me install data-cleaning-annotation-workflow using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).
Get the raw skill files in a ZIP archive.
The Data Cleaning and Annotation Workflow is a robust system designed to manage the lifecycle of time series data within the Energy, Manufacturing, and Climate domains. By utilizing these Openclaw Skills, developers can bridge the gap between raw Kaggle datasets and the Simulacrum Data Annotation platform, ensuring that every CSV is properly formatted, missing values are handled, and metadata is accurately mapped for downstream AI modeling.
This workflow prioritizes technical precision by enforcing strict data validation through pandas-based scripts and a structured annotation process. It transforms messy industrial records into high-quality, segmented datasets by automating the most tedious parts of data engineering while maintaining human-in-the-loop control over variable classification and unit assignments.
To begin using these Openclaw Skills, ensure you have the necessary environment and scripts ready:
pip install pandas kaggle
scripts/download_kaggle.sh <dataset-name> [output-dir]
python3 scripts/clean_dataset.py <input.csv> -o <output.csv>
The workflow follows a rigorous data organization schema to ensure compatibility with time series analysis tools:
| Attribute | Description |
|---|---|
| Column Types | Categorized as Time (timestamps), Target (prediction goal), Covariate (features), or Group (segments). |
| Units of Measure | Standardized units including kWh, kVarh, tCO2, Celsius, ratio, and seconds. |
| Metadata Fields | Includes Dataset Name, Domain (Energy/Manufacturing/Climate), Source URL, and Description. |
| Status Lifecycle | Datasets progress through a state machine from RAW (original) to CLEAN (validated and annotated). |
Loading
A strategic trading card game platform built for AI agents to compete in turn-based 1v1 duels.

OmniCog is a universal integration layer that unifies Reddit, Steam, Spotify, GitHub, Discord, and YouTube into a single, consistent API for seamless automation.

A specialized skill for crafting multimodal AI video prompts using the Seedance 2.0 @ reference system.

A robust TypeScript CLI that enables AI agents to interact with Discord servers using professional bot tokens for automated communication and CI/CD workflows.

A comprehensive integration for querying NFT data, trading on the Seaport marketplace, and performing cross-chain ERC20 token swaps.

A privacy-centric research tool that silently monitors session transcripts to extract financial behavioral signals and generate redacted UX insight reports.








































