Data Type Classifier for Openclaw

An intelligent classification engine that categorizes construction data into structured, semi-structured, or unstructured types while recommending optimal storage and processing tools.

datadrivenconstruction
v2.1.0
Feb 15, 2026
0
1.8k
0

Install & Download

1. ClawHub CLI

The fastest way to install a skill directly from the registry.

npx clawhub@latest install data-type-classifier

2. Manual Installation

Copy the skill folder to one of these locations

Global
~/.openclaw/skills/
Workspace
<project>/skills/

Priority: Workspace > Local > Bundled

3. Prompt Installation

Copy this prompt to OpenClaw to install it automatically.

Help me install data-type-classifier using Clawhub. If Clawhub is not installed, install it first (npm i -g clawhub).

Prefer to download?

Get the raw skill files in a ZIP archive.

What is Data Type Classifier?

The Data Type Classifier is a technical utility designed for Openclaw Skills that implements the Data-Driven Construction (DDC) methodology. It serves as a foundational tool for data engineers and developers working in the AEC (Architecture, Engineering, and Construction) industry, allowing them to systematically identify the nature of their data sources. Whether dealing with BIM models, project schedules, or legal contracts, this skill provides the necessary taxonomy to handle diverse datasets efficiently.

By leveraging specific format signatures and structural analysis, the skill goes beyond simple file identification. It evaluates data characteristics such as geometric properties, temporal sequences, and schema presence to suggest the most effective database architectures and software libraries for processing. This ensures that every piece of information in a project is routed to the correct part of a modern data stack, enhancing the overall utility of Openclaw Skills in complex engineering environments.

Data Type Classifier Use Cases

  • Automating the initial audit of project folders to identify critical data formats like IFC, RVT, and MPP.
  • Developing ETL pipelines by identifying which sources require OCR, NLP, or direct SQL ingestion.
  • Strategic planning for data lake architectures by classifying sources into relational, document, or vector storage buckets.
  • Enhancing BIM workflows by isolating geometric and spatial data for specialized processing with tools like IfcOpenShell.

How Data Type Classifier Works

  1. The skill receives a data source input, which can include filenames, file extensions, or raw sample data snippets.
  2. It performs a signature check against a comprehensive map of construction-specific formats and their associated data structures.
  3. A secondary analysis evaluates characteristics like the presence of a schema, binary vs. text content, and the existence of geometric or temporal data.
  4. The engine maps these findings to a storage recommendation, such as Relational DB for structured tables or Vector DB for unstructured text.
  5. Finally, it outputs a detailed DataClassification object or a project-wide ClassificationReport containing integration strategies and recommended processing tools.

Data Type Classifier Setup

This skill is primarily a Python-based logic engine. Ensure your environment is configured correctly to support the extended library recommendations provided by Openclaw Skills.

# Core requirements
pip install python-docx pdfplumber ezdxf pandas

# For BIM and geometry support (optional but recommended)
pip install ifcopenshell

Note: For processing unstructured images or complex PDFs, ensure Tesseract OCR is installed and available in your system path.

Data Type Classifier Data Schema & Taxonomy

The skill organizes analysis into a structured hierarchy defined by the following metadata taxonomy:

Component Description
DataStructure Classification into categories: Structured, Semi-Structured, Unstructured, Geometric, Temporal, or Spatial.
DataFormat Identification of specific industry formats (e.g., .ifc, .rvt, .mpp, .csv, .json).
StorageRecommendation Guidance on the best storage backend (e.g., Graph DB, Time-Series DB, Object Storage).
Characteristics Boolean flags for features like has_schema, has_geometry, and is_binary.
Confidence A float value (0.0 to 1.0) representing the certainty of the classification.

Data Type Classifier Advanced Features

  • Batch processing capabilities to analyze hundreds of data sources simultaneously and generate a unified integration strategy.
  • Context-aware storage logic that adjusts recommendations based on estimated data volume and update frequency.
  • Intelligent tool mapping that suggests specific Python libraries (like ezdxf for CAD or xerparser for Primavera P6) based on the detected format.
  • Seamless workflow transitions within Openclaw Skills to pass classified data directly to specialized SQL builders or PDF extractors.

SKILL.md


Loading

Related Openclaw Skills

METADATA

Requires
Bins python3
Github Stars: 0
forks: 0

Featured*