• bignose@programming.dev
    link
    fedilink
    English
    arrow-up
    1
    ·
    5 days ago

    Copyright grounds its justification in the creative expression of people. By creative transformation, new works can be created and some of those are deemed not restricted by copyright in the earlier work.

    An LLM is not a person. An LLM has no mind and no creative expression. The training of an LLM and the inference from the LLM are, by definition, non-creative mechanical transformation from the input training data.

    The copyright of the input work should, by this logic, survive whole in the output from that mechanical non-creative process. That is substantially different from the process by which humans learn and create.

    • encelado748@feddit.org
      link
      fedilink
      arrow-up
      1
      ·
      5 days ago

      You can make the case that an LLM is not a person, you cannot make the case that LLM work is not creative. The grounding justification of copyright lacked a foundational example of a non human creative process. If each single token emitted by the LLM is the result of math on the entire knowledge matrix, then each single token is derived from all the copyrighted documents used in training in a percentage you cannot even estimate. The mechanical transformation of an LLM is unknowable, unquantifiable and non deterministic. The copyright claim makes no sense in this case. A new shared mechanism must be put into place if we want to translate copyright through such mechanism. Lot of the training data used by LLM is synthetic so it pass through multiple layers of unknowable, unquantifiable, non deterministic “mechanisms”.