LLM for Coding Benchmarks and Datasets
LiveCodeBench, SWEBench, Aider Polyglot, BBH, HumanEval, MBPP, Common Crawl (Time Span, Dataset Size, Data Format, LLM Testing Capability)

Search for a command to run...
Articles tagged with #llm
LiveCodeBench, SWEBench, Aider Polyglot, BBH, HumanEval, MBPP, Common Crawl (Time Span, Dataset Size, Data Format, LLM Testing Capability)

Turning a base model into a reasoning model is essentially a post-training + data problem

If LLM is compared to a person, then the early training is like general education, and the later training is like vocational training.

Basic RAG Patterns Naive RAG Simple retrieve-then-generate pipeline Direct semantic search → context injection → LLM generation Works well for straightforward Q&A over documents Advanced RAG Pre-retrieval query optimization (query rewriting, expan...

Comprehensive analysis of the most advanced coding-focused LLM models: Architecture, Context Window, Cost, Performance & Practical Coding Capabilities

Imagine that you have a pre-trained model which works well on task A. Now you get some similar tasks called task B, C, D. How to fine-tune your base model for those new tasks? Can we just update partial of the parameters for each task and run those t...
