Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
36.0
TFLOPS
AlexK
aiDev03312
3
3
Follow
etemiz's profile picture
1 follower
ยท
2 following
AI & ML interests
None yet
Recent Activity
replied
to
etemiz
's
post
about 2 hours ago
maybe one of the hard parts of my kind of fine tuning is evals. where do you want model to go? how do you find true answer of a hardly debated issue? instead of manually writing answers in many domains (which i can't do, i don't know many answers in many domains) i rely on a few tricks: - find other aligned llms and get ideas from them - rank many llms in AHA leaderboard and get ideas from top ones also rejecting the worst ones - do mixture of agents of the above to get a collective answer - and lately, compare the answers coming from my fine tunes with base models, assume my fine tune is preferred if there is a difference these still can't find perfect answers but they kick the model in the right direction. and that may be a big deal. can't claim my fine tune (ostrich) knows every truth. but it may make more sense to claim if the base differs with fine tune most probably fine tune is the better answer. making better evals ends up training better models. and produce better AHA ranking. which further sharpens evals. this feedback loop is going to be useful for a while.
reacted
to
etemiz
's
post
with ๐
about 2 hours ago
maybe one of the hard parts of my kind of fine tuning is evals. where do you want model to go? how do you find true answer of a hardly debated issue? instead of manually writing answers in many domains (which i can't do, i don't know many answers in many domains) i rely on a few tricks: - find other aligned llms and get ideas from them - rank many llms in AHA leaderboard and get ideas from top ones also rejecting the worst ones - do mixture of agents of the above to get a collective answer - and lately, compare the answers coming from my fine tunes with base models, assume my fine tune is preferred if there is a difference these still can't find perfect answers but they kick the model in the right direction. and that may be a big deal. can't claim my fine tune (ostrich) knows every truth. but it may make more sense to claim if the base differs with fine tune most probably fine tune is the better answer. making better evals ends up training better models. and produce better AHA ranking. which further sharpens evals. this feedback loop is going to be useful for a while.
replied
to
etemiz
's
post
3 days ago
getting ready to fine tune 3.8 - evolutionary strategies - behavior steering experiments - expanded dataset - bringing back ORPO - more orthogonal evals to keep overfitting minimum - most probably will take abliterations as base, either mine or somebody else's - random entropy addition from huggingface fine tunes (take what is popular on hf and randomly introduce into the lineage) - bring more vibe coding: turns out LLMs know how to fine tune
View all activity
Organizations
None yet
aiDev03312
's datasets
None public yet