public
Beijing-AISI/panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.
public
IDEAS-Lab-Northwestern/datagen-clutter-v1-joint-5cam
datagen-clutter-v1-joint-5cam
Auto-generated SFT dataset for the clutter (pick-out-of-clutter → place-in-goal) family — a
cuRobo-planned, physics- & LTL-safety-checked demonstration set, already converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — collected on all 55 clutter_pickup base tasks.
Per task: 40 success + LTL-safe trajectories → 2,200 episodes total.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-clutter-v1-joint-5cam.
public
icdn11/content-202606086b69
No description provided.
public
RalphLabsAI/proof-bundles
No description provided.
public
chenzeyang1/datasets
No description provided.
public
Junrui1202/zhoblimp
ZhoBLIMP Dataset
This is the Chinese version of the BLIMP (Benchmark of Linguistic Minimal Pairs) dataset.
Dataset Structure
The dataset contains 118 subtasks, each testing different linguistic phenomena in Chinese.
Usage
from datasets import load_dataset
# Load a specific subtask
dataset = load_dataset("your_username/zhoblimp", "BA_verb_le_b")
# Load all configs
all_configs = load_dataset("your_username/zhoblimp", "all")
Subtasks… See the full description on the dataset page: https://huggingface.co/datasets/Junrui1202/zhoblimp.
public
icdn18/content-20260609b21c
No description provided.
public
EER6/nvidia-OpenCodeInstruct-refined
nvidia-OpenCodeInstruct-refined
A strictly quality-filtered subset of nvidia/OpenCodeInstruct (5M examples). This is a strict subset of EER6/nvidia-OpenCodeInstruct-broad.
Filtering criteria
Both conditions must be satisfied:
Criterion
Threshold
LLM judge min score
= 5 (out of 5)
Unit test pass rate (average_test_score)
= 1.0
LLM judge min score is the minimum across all three dimensions in the llm_judgement field:
requirement_conformance — does… See the full description on the dataset page: https://huggingface.co/datasets/EER6/nvidia-OpenCodeInstruct-refined.
public
vincewin/CREST_data
CREST forcing (parquet)
EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year.
dir
variable
source
cadence
mrms/
precipitation
MRMS QPE (corrected)
hourly
temp/
2 m temperature
NLDAS-2 FORA
hourly
pet/
potential ET
FEWS NET daily PET
daily
Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the
EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI.
Download + extract one year, e.g.:
from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.
public
lasrprobegen/sycophancy-activations
No description provided.
public
Weiyun1025/InternVL-Performance
No description provided.
public
ESA-philab/S1_Uncropped
No description provided.
public
leduytho/Egoverse_videos
No description provided.
public
notamitgamer/usercontent
notamitgamer/bsc
BSc Computer Science Honours (WBSU)
This repository serves as a comprehensive, live archive of my 4-year academic journey in Computer Science at Acharya Prafulla Chandra College (APC). It contains practical implementations, assignments, and study materials following the WBSU curriculum.
Quick Links
GitHub Repo: @notamitgamer/bsc
Web View notamitgamer.github.io/bsc
Web View: code.amit.is-a.dev
Live Portfolio: amit.is-a.dev
GitHub… See the full description on the dataset page: https://huggingface.co/datasets/notamitgamer/usercontent.
public
tasksource/bigbench
BIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version.
dataset = load_dataset("tasksource/bigbench",'movie_recommendation')
Code to reproduce:
https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing
Datasets are capped to 50k examples to keep things light.
I also removed the default split when train was available also to save space, as default=train+val.
@article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bigbench.
public
icdn14/content-202605255781
No description provided.
public
icdn15/content-20260609b294
No description provided.
public
icdn14/content-2026052887f2
No description provided.
public
pdfqa/pdfQA-Benchmark
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
The dataset is organized to support:
Raw document processing research
Structured extraction pipelines
Retrieval-augmented QA
End-to-end document reasoning systems
It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.
public
phamthithu1993/phamthithu1993
No description provided.
public
vibrantlabsai/fiqa
FiQA Dataset for RAG Evaluation
The FiQA (Financial Opinion Mining and Question Answering) dataset reformatted specifically for evaluating Retrieval-Augmented Generation (RAG) systems. This dataset contains financial domain questions with ground truth answers and retrieved contexts, making it ideal for testing RAG pipelines on domain-specific content.
Recommended Usage: ragas_eval_v3
The ragas_eval_v3 configuration is the primary and recommended way to use this… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/fiqa.