Genomic Sequence Benchmark Health Archive
Category: Health License: cc-by-nd-4.0 Downloads: 717 Rows: 2375 CSV
Dataset card Data preview Source License Citation

Dataset Card

Dataset Description

Genomic Sequence Benchmark Health Archive contains structured health-related records designed for research, analytics, dashboard development, and machine learning experiments. The dataset includes realistic patient, clinical, administrative, and observational variables with category-specific value distributions and no repeated column pattern.

Dataset Summary

This dataset provides structured CSV data for analysis, reporting, research, dashboards and machine learning workflows.

Dataset Structure

The file is available in CSV format and can be used with Python, R, Excel, Google Sheets, Power BI or other data analysis tools.

Data Preview

sequence_id gene_region sequence_length gc_content variant_type risk_marker population_group analysis_batch
1 promoter 498 0.538 Insertion Moderate Group C BATCH-207
2 intronic 189 0.674 SNP None Group D BATCH-340
3 promoter 226 0.518 Deletion Elevated Group A BATCH-441
4 intergenic 265 0.357 CNV Elevated Group C BATCH-781
5 promoter 292 0.459 SNP None Group B BATCH-729
6 exonic 299 0.645 SNP None Group B BATCH-264
7 promoter 350 0.446 Insertion Elevated Group C BATCH-153
8 promoter 135 0.338 Insertion Low Group D BATCH-694
9 intronic 493 0.567 Insertion None Group B BATCH-813
10 promoter 112 0.518 SNP Elevated Group D BATCH-183

License

cc-by-nd-4.0

Citation

Genomic Sequence Benchmark Health Archive. Health Research Data Repository, Version 2.6, 2018.