users lookin at my profile counter:
(its mostly myself)

heyhi ~ i make posts and hope that peeps comment on them so i can comment on them and kinda have a conversation sorta.

come and have a tea with me 🍵 <3

  • 70 Posts
  • 731 Comments
Joined 3 years ago
cake
Cake day: July 5th, 2023

help-circle



  • yeeesssss the playground truly was a playground with the purpose of… play.

    but… don’t tell this to anyone… but

    secret

    you can have fun with local language models and still have transitioned.

    these smaller local models tool a massively smaller footprint to train, are much less restrictive and got a generally different vibe to them.
    they are dumber than what big evil tech is selling to us, but… they are locally running on your lappytoppy or desktop.

    for an easy gateway, pull ollama and run ollama run qwen3.5:4b. that commands pulls and runs a pretty okay Qwen model on your GPU or CPU (whatever ollama finds on your machine).
    (qwen is a chinese model, by alibaba, so… if you don’t like china, maybe run gemma4:e4b instead. it’s by google…)

    if you want more of those “literally just a text predictor” vibes, i hiiiiighly recommend llama-cpp. it’s faster, it’s good, it’s what everything else is built on AND it even has a nice ui.
    you can pull the release here (i recommend the “vulkan” one, even if you only have an iGPU).
    once unpacked, pull some “gguf” model that fits in your VRAM or RAM from somewhere. if you have about 15 GB of RAM free on your system, you may want to try Qwen3.6-35B-A3B or it’s no-conversation-training counterpart Qwen3.5-35B-A3B-Base at some lower quantization like the “IQ2_M” one.

    • click on the quant
    • click the “download” that appears in the sidebar
    • wait for it to finish

    finally, run it with the “llama-server” binary you find in the llama cpp folder like so:

    # ctk and ctv make it go faster
    ./llama-server --model ~/.llama-cpp/models/qwen3-coder/Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf -ctk q4_0 -ctv q4_0 --port 9090
    0.00.036.875 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
    0.00.037.353 W srv  llama_server: -----------------
    0.00.037.356 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
    0.00.037.356 W srv  llama_server: this can be a security risk (cross-origin attacks)
    0.00.037.356 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
    0.00.037.356 W srv  llama_server: -----------------
    0.00.038.704 I srv    load_model: loading model '/home/maria/.llama-cpp/models/qwen3-coder/Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf'
    0.00.641.956 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
    0.12.362.730 I cmn          init: llama threadpool init, n_threads = 2
    0.13.447.376 I srv    load_model: initializing, n_slots = 4, n_ctx_slot = 101632, kv_unified = 'true'
    0.13.474.204 I srv  llama_server: model loaded
    0.13.474.220 I srv  llama_server: listening on http://127.0.0.1:9090/
    

    open that page http://127.0.0.1:9090/ and see what goes.

    ollama automatically creates a server in the background which you can curl from like this:

    curl 127.0.0.1:11434/v1/chat/completions -d '{
    > "messages": [
    >   { "role": "user", "content": "heyhiiii ~ ~ ~ ~ ~ how we doinnnn?" }
    > ]
    > }'
    

    llama-cpp is the same, but you gotta start the server like i showed earlier. these are standard “OpenAI-compatible” APIs. meaning: you can use them in agents and such if you wanna.

    buuuuuut yeaaaa it kinda still is very much… LM stuff. so i can definitely understand staying far away from this.



  • is this referring to me specifically or your perspective on the matter? >v<

    i feel like having someone feed you berries and fruits is not on par with intercourse. it sits well above that.

    having someone you feel so safe with, and who feels so safe with you that they are willing to do that, unprompted! that’s about as strong as it gets.
    imagine placing a strawberry oh so tenderly against someones lips and have them first pursed, but then soften and open up for you. imagine lightly sliding it in and then have them close up again, pursing their lips a bit, chewing and swallowing.

    and finally feel a bit hesitant to want more.

    i feel that this is super lovely and sweet and so very full of feelings one would not expect to find in such a situation.












  • maria ~@lemmy.blahaj.zoneto196@lemmy.blahaj.zonerule
    link
    fedilink
    English
    arrow-up
    9
    ·
    2 days ago

    but that is terrible!!! ignoring them… no, nonono, I couldnt do that. Neglegance goes beyond punishment, that is just… cruel. in the sad way. in the lonely cold not-fun-al-all way.

    thats no fun. not for me, not for the subbie, no I don’t want that, thats just so very sad I couldnt bare neglecting… no, no that feels bad… imma stay with teasing




  • maria ~@lemmy.blahaj.zoneto196@lemmy.blahaj.zonerule
    link
    fedilink
    English
    arrow-up
    17
    ·
    2 days ago

    hmmm in my mind rewards are like - overwhelming your subbie and really going ham on them and kissing them all over furiously and such, making them lose that feeling of needing to hide how nice they feel. but a punishment! >o< im honestly not sure what that would even be… iguess teasing. hurting them feels wrong…