• leanleft@lemmy.ml
    link
    fedilink
    English
    arrow-up
    4
    ·
    edit-2
    1 day ago

    according to performance on standard benchmark. somewhat covered by the controversy surrounding the term: benchmaxing.
    if you see all benefit as a linear one dimensional height on a bar graph…
    its almost like you assume that the previous model gave the same exact answer(same style) and the new mode gave the same exact answer PLUS additional useful information.
    it might be convenient if measuring progress was so simple. but unfortunately/fortunately , it is not so simple . the most important benchmark are the comparison of outcomes on the problems that YOU have & prompts that YOU can(will) write. nothing else matters for YOU.

    • i admit benchmarks are well designed to objectively measure competence on challenging problems that require skill and really need only ONE correct answer.
    • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
      link
      fedilink
      arrow-up
      6
      ·
      1 day ago

      Sure, a benchmark doesn’t capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn’t really about the nuance, it’s showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.

      • leanleft@lemmy.ml
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 hours ago

        i see what your saying. i didnt mean to discredit standard benchmarks entirely.
        i guess its obvious that it measures capability regardless of imprecision.
        2 major proposed changes:
        **first, i dont really know. aside from saying “benchmark your own prompt+usecase”
        a proposed plan:

        • approach one: pay attention and credit new or improved architecture designs and research.
        • approach two: spend more attention on benchmarks. especially specific benchmarks ( that are not focused with industrial domain tasks.) **domain task pursuit, is useful!.. but it depends on if your interest align to popular domains.
        • approach three: if willing to utilize remotely hosted models. rating should also take in consideration… tools and everything else: websearch performance, RAG performance, smooth interface, pref/balance between speed vs comprehensiveness, cost (if relevant), etc… .
        • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
          link
          fedilink
          arrow-up
          3
          ·
          4 hours ago

          Honestly, I think the most reasonable approach is just to see what other people’s experience is like and which models are well regarded, then try them out and see which one is the best fit for what you’re doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.