Data Cleaning and Annotation Workflow for Openclaw

A comprehensive end-to-end pipeline for discovering, cleaning, and annotating industrial time series datasets for machine learning.

deyashmukh
v1.0.0
Feb 18, 2026
0
2k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install data-cleaning-annotation-workflow

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install data-cleaning-annotation-workflow using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Data Cleaning and Annotation Workflow?

The Data Cleaning and Annotation Workflow is a robust system designed to manage the lifecycle of time series data within the Energy, Manufacturing, and Climate domains. By utilizing these Openclaw Skills, developers can bridge the gap between raw Kaggle datasets and the Simulacrum Data Annotation platform, ensuring that every CSV is properly formatted, missing values are handled, and metadata is accurately mapped for downstream AI modeling.

This workflow prioritizes technical precision by enforcing strict data validation through pandas-based scripts and a structured annotation process. It transforms messy industrial records into high-quality, segmented datasets by automating the most tedious parts of data engineering while maintaining human-in-the-loop control over variable classification and unit assignments.

Data Cleaning and Annotation Workflow Use Cases

  • Preparing steel industry energy consumption data for demand forecasting.
  • Cleaning manufacturing emission datasets to identify carbon footprint patterns.
  • Standardizing climate monitoring data for environmental impact analysis.
  • Moving raw Kaggle CSV files through a professional annotation pipeline to reach verified CLEAN status.

How Data Cleaning and Annotation Workflow Works

  1. Search for and retrieve specific time series datasets from Kaggle using automated Openclaw Skills or the web interface.
  2. Execute a localized Python script to perform automated data cleaning, including duplicate removal and median imputation.
  3. Upload the raw dataset to the annotation platform with essential source metadata and domain classification.
  4. Define column headers by selecting appropriate types such as Time, Target, Covariate, or Group.
  5. Map physical units like kWh, tCO2, or ratios to ensure mathematical consistency across the dataset.
  6. Bulk-assign group tags to variables to enable multi-segment data analysis.
  7. Process the final upload to transition the dataset from RAW to CLEAN status.

Data Cleaning and Annotation Workflow Setup

To begin using these Openclaw Skills, ensure you have the necessary environment and scripts ready:

  1. Install dependencies:
pip install pandas kaggle
  1. Download a dataset via the Kaggle CLI:
scripts/download_kaggle.sh <dataset-name> [output-dir]
  1. Run the cleaning script to prepare your data:
python3 scripts/clean_dataset.py <input.csv> -o <output.csv>

Data Cleaning and Annotation Workflow Data Schema & Taxonomy

The workflow follows a rigorous data organization schema to ensure compatibility with time series analysis tools:

Attribute Description
Column Types Categorized as Time (timestamps), Target (prediction goal), Covariate (features), or Group (segments).
Units of Measure Standardized units including kWh, kVarh, tCO2, Celsius, ratio, and seconds.
Metadata Fields Includes Dataset Name, Domain (Energy/Manufacturing/Climate), Source URL, and Description.
Status Lifecycle Datasets progress through a state machine from RAW (original) to CLEAN (validated and annotated).

Data Cleaning and Annotation Workflow Advanced Features

  • Bulk configuration tools to apply metadata across hundreds of columns simultaneously.
  • Automated timestamp conversion that handles various datetime formats found in public datasets.
  • Intelligent group tagging that links target variables and covariates for complex segmented training.
  • Integrated Kaggle CLI support for programmatic dataset acquisition as part of larger Openclaw Skills automation.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Github Stars: 0
forks: 0

Featured*