GaiaEx AcademyGaiaEx Academy
ന്യൂറൽ നെറ്റ്‌വർക്കുകൾ: പെർസെപ്ട്രോണുകൾ മുതൽ ഡീപ് ലേണിംഗ് വരെ
ഡെവലപ്പർAI & ML12 min read

ന്യൂറൽ നെറ്റ്‌വർക്കുകൾ: പെർസെപ്ട്രോണുകൾ മുതൽ ഡീപ് ലേണിംഗ് വരെ

ലളിതമായ ഫങ്ഷനുകളുടെ ലെയറുകൾ സങ്കീർണ്ണമായ പാറ്റേണുകൾ എങ്ങനെ പഠിക്കുന്നു

പോസ്റ്റുകൾ പങ്കിടുക

ആരും മനസ്സിലാക്കാത്ത Move

2016 മാർച്ച് 10-ന്, Seoul-ൽ അഞ്ച്-match series-ന്റെ രണ്ടാം game-ൽ, AlphaGo എന്ന ഒരു machine board-ന്റെ അഞ്ചാം line-ൽ ഒരു black stone വച്ചു. Move ഇത്ര strange ആയിരുന്നു, human commentator-മാർ അത് ഒരു mistake ആണ് എന്ന് കരുതി. ഒരു professional ഇത് play ചെയ്താൽ ഒരു beginner-നെ scold ചെയ്യും എന്ന് ഉച്ചത്തിൽ പറഞ്ഞു. Lee Sedol — Go game-ന്റെ 18-time world champion — collect ചെയ്യാൻ room-ൽ നിന്ന് നിന്ന് പോയി.

അത് ഒരു mistake ആയിരുന്നില്ല. അമ്പത് move-കൾക്ക് ശേഷം, ആ stone — ഇപ്പോൾ "Move 37" ആയി famous ആണ് — മുഴുവൻ game-ഉം turn ചെയ്ത hinge ആയിരുന്നു. AlphaGo ജയിച്ചു. Lee Sedol, ജീവിക്കുന്ന ഏറ്റവും great player-കളിൽ ഒരാളായി consider ചെയ്യപ്പെടുന്നു, series 4–1 നഷ്ടപ്പെട്ടു. ഇവിടെ പ്രധാനമായ ഭാഗം: ഒരു മനുഷ്യനും AlphaGo-യെ ആ move പഠിപ്പിച്ചില്ല. ഒരു textbook-ലും ഇത് ഉണ്ടായിരുന്നില്ല. Machine അതിനെതിരെ millions games play ചെയ്ത്, ഒരു pattern emerge ചെയ്യുന്നത് വരെ billions ചെറിയ internal numbers adjust ചെയ്ത്, സ്വയം discover ചെയ്തതാണ്.

Go-ക്ക് observable universe-ലെ atom-കളെക്കാൾ കൂടുതൽ possible board position-കൾ ഉണ്ട്. നിങ്ങൾക്ക് അത് brute-force ചെയ്യാൻ കഴിയില്ല. അതിന് rules എഴുതാൻ കഴിയില്ല. Superhuman level-ൽ അത് play ചെയ്യാനുള്ള ഏകവഴി learn ചെയ്യുകയാണ് — learning ചെയ്യുന്ന കാര്യം ഒരു neural network ആയിരുന്നു: mathematical unit-കളുടെ ഒരു stack, മതിയായ ഡേറ്റയും മതിയായ adjustment-ഉം ഉണ്ടെങ്കിൽ, ഒരു മനുഷ്യനും fully articulate ചെയ്യാൻ കഴിയാത്ത structure കണ്ടെത്താൻ കഴിയും.

നിങ്ങൾ talk ചെയ്യുന്ന language model-നും, fraudulent card transaction-കൾ flag ചെയ്യുന്ന system-നും, quantitative fund-കൾ ഇപ്പോൾ financial market-കൾ-ലേക്ക് point ചെയ്യുന്ന model-കൾക്കും power നൽകുന്നത് ഇതേ machinery ആണ് — അതേ weighted sum-കൾ, അതേ training loop, അതേ gradient nudge-കൾ. ആ box-നുള്ളിൽ യഥാർത്ഥത്തിൽ എന്ത് സംഭവിക്കുന്നു എന്നതിനെക്കുറിച്ചാണ് ഈ ലെസൺ.

Perceptron: അതെല്ലാം എവിടെയാണ് ആരംഭിക്കുന്നത്

ഒരു simple spam filter മുതൽ GPT വരെ AlphaGo വരെയുള്ള ഓരോ neural network-ഉം ഒരു single idea-യിൽ നിന്ന് descend ചെയ്യുന്നു: 1958-ൽ Frank Rosenblatt invent ചെയ്ത perceptron. ഇത് ഒരു biological neuron-ൽ loosely inspire ചെയ്ത ഒരു mathematical model ആണ്. ഒരു real neuron അതിന്റെ dendrite-കൾ വഴി electrical signal-കൾ receive ചെയ്യുന്നു, cell body-യിൽ അവ accumulate ചെയ്യുന്നു, combined input ഒരു threshold cross ചെയ്യുന്നു എങ്കിൽ മാത്രം അതിന്റെ axon-ലൂടെ ഒരു signal fire ചെയ്യുന്നു. Perceptron voltage-ന് പകരം numbers ഉപയോഗിച്ച് അതേ കാര്യം ചെയ്യുന്നു.

ഒരു perceptron input-കളുടെ ഒരു vector x എടുക്കുന്നു, ഓരോ input-നും ഒരു learned weight w കൊണ്ട് multiply ചെയ്യുന്നു, products add ചെയ്യുന്നു, threshold shift ചെയ്യുന്ന ഒരു bias term b add ചെയ്യുന്നു, ഫലം ഒരു activation function-ലൂടെ pass ചെയ്യുന്നു. ഫലം bar clear ചെയ്താൽ, neuron "fire" ചെയ്യുന്നു; അല്ലാത്തപക്ഷം quiet ആയി തുടരുന്നു. ഇതെല്ലാം ഒരു compact expression-ൽ: output = f(w·x + b).

Weight-കളെ volume knob-കൾ ആയി ചിന്തിക്കുക. ഓരോ input-ഉം ചില loudness-ഓടെ എത്തുന്നു, weight അത് ഈ particular neuron-ന് എത്ര പ്രധാനമാണ് എന്ന് തീരുമാനിക്കുന്നു. ഒരു fraudulent trade spot ചെയ്യാൻ learn ചെയ്യുന്ന ഒരു neuron "transaction size relative to account history"-ന്റെ weight crank up ചെയ്യുകയും "time of day"-ന്റെ weight turn down ചെയ്യുകയും ചെയ്യാം. Learning, its core-ൽ, ആ knob-കൾ ശരിയായ settings-ലേക്ക് turn ചെയ്യുന്നത് മാത്രമാണ്, glamorous ഒന്നും അല്ല.

ഒരു single perceptron-ന് ഒരു hard limit ഉണ്ട്: ഇതിന് ഡേറ്റ ഒരു straight line കൊണ്ട് മാത്രം separate ചെയ്യാൻ കഴിയും (അല്ലെങ്കിൽ, higher dimension-കളിൽ, ഒരു flat hyperplane). "ഈ point line-ന് മുകളിലാണോ താഴെയാണോ?" എന്നതിന് ഉത്തരം പറയാം, പക്ഷേ "ഈ point curved region-ന് ഉള്ളിലാണോ?" എന്നതിന് അല്ല. Historic example XOR problem ആണ് — ഒരു single perceptron-ന് provably learn ചെയ്യാൻ കഴിയാത്ത ഒരു pattern, 1970-കളിൽ AI funding-നെ years-ന് freeze ചെയ്ത ഒരു limitation. Hindsight-ൽ escape almost absurdly simple ആണ്: നിരവധി perceptron-കൾ layer-കളാക്കി stack ചെയ്യുക, അവയ്ക്കിടയിൽ ഒരു nonlinear activation insert ചെയ്യുക. അത് ചെയ്യുക, resulting network space-നെ arbitrarily complex shapes-ലേക്ക് bend ചെയ്യാം, fold ചെയ്യാം, carve ചെയ്യാം.

പ്രധാന insight: ഒരു neural network magic അല്ല, ഒരു brain-ഉം അല്ല. അത് dot product-കളുടെയും simple nonlinear function-കളുടെയും ഒരു tower ആണ്, ദശലക്ഷക്കണക്കിന് weight-കൾ better values-ലേക്ക് nudge ചെയ്യപ്പെടുന്നു. "Intelligence" പൂർണ്ണമായി ആ weight-കൾ ഏത് numbers-ൽ settle ചെയ്യുന്നു എന്നതിലാണ് ജീവിക്കുന്നത് — ആ numbers ഡേറ്റയാൽ കണ്ടെത്തുന്നു, ഒരു programmer എഴുതുന്നതല്ല.
Perceptron: weighted sum + nonlinearity x₁ x₂ w₁ w₂ Σ + b linear part σ, ReLU… a ŷ Stack layers + nonlinear activations → universal approximator (in principle)
ഒരു neuron: dot product, bias, പിന്നീട് activation — ഓരോ deep net-ന്റെയും atom.

Layer-കൾ, Depth, 'Deep' Learning എന്തുകൊണ്ട്

ഒരു neuron ഒരു knob ആണ്. ഒരു useful network ആയിരക്കണക്കിന് knob-കൾ ആണ്, information forward pass ചെയ്യുന്ന layer-കളായി organize ചെയ്ത, ഒരു bucket brigade-നെപ്പോലെ.

  • Input layer — നിങ്ങളുടെ raw feature-കൾ enter ചെയ്യുന്നിടം: ഒരു image-ന്റെ pixel-കൾ, ഒരു sentence-ന്റെ words, അല്ലെങ്കിൽ ഒരു trading model-ന്, price, volume, volatility, order-book imbalance പോലുള്ള കാര്യങ്ങൾ.
  • Hidden layer-കൾ — ഇടയിലുള്ള layer-കൾ, real work നടക്കുന്നിടം. ഓരോ layer-ഉം previous layer-ന്റെ output-കൾ എടുത്ത് പുതിയ, more abstract feature-കളാക്കി recombine ചെയ്യുന്നു.
  • Output layer — final answer: ഒരു probability, ഒരു price forecast, ഒരു class label.

Deep learning-ലെ "deep" എന്ന വാക്ക് ജസ്റ്റ് നിരവധി hidden layer-കൾ എന്നാണ് അർത്ഥമാക്കുന്നത്. Depth ഒരു നിർദ്ദിഷ്ട കാര്യം നിങ്ങൾക്ക് buy ചെയ്യുന്നു: feature-കളുടെ ഒരു hierarchy. ഒരു image network-ൽ, first layer edge-കൾ detect ചെയ്യുന്നു, next layer shapes-ലേക്ക് edge-കൾ assemble ചെയ്യുന്നു, next layer object-കളിലേക്ക് shape-കൾ assemble ചെയ്യുന്നു. ആരും "edge detector" അല്ലെങ്കിൽ "eye detector" programme ചെയ്തിട്ടില്ല — ആ concept-കൾ training-ൽ നിന്ന് emerge ചെയ്യുന്നു. Network അതിന്റെ സ്വന്തം intermediate vocabulary invent ചെയ്യുന്നു.

ഇവിടെ ഒരു beautiful theoretical result ഉണ്ട്, universal approximation theorem: മതിയായ neuron-കൾ ഉള്ള ഒരു single hidden layer ഉള്ള ഒരു network-ന് ഏത് continuous function-ഉം arbitrary accuracy-ലേക്ക് approximate ചെയ്യാം. മറ്റ് വാക്കുകളിൽ, ശരിയായ network-ന് നിങ്ങളുടെ ഡേറ്റയിൽ exist ചെയ്യുന്ന ഏത് pattern-ഉം theory-യിൽ represent ചെയ്യാം. Catch — ഒരു വലിയ catch ആണ് — "in principle" heavy lifting ചെയ്യുകയാണ്. Theorem promise ചെയ്യുന്നത് അത്തരം ഒരു network exist ചെയ്യുന്നു എന്നാണ്; നിങ്ങൾക്ക് അത് find ചെയ്യാൻ കഴിയുമോ, അത് pin down ചെയ്യാൻ മതിയായ ഡേറ്റ നിങ്ങൾക്ക് ഉണ്ടോ, നിങ്ങൾ chase ചെയ്യുന്ന pattern real ആണോ എന്നൊന്നും അത് ഒന്നും പറയുന്നില്ല. ധനകാര്യം പോലുള്ള low-signal domain-കളിൽ, "representable"-ഉം "learnable"-ഉം തമ്മിലുള്ള gap കൃത്യമായി മിക്ക model-കളും die ചെയ്യുന്നിടത്താണ്.

Depth effort-ന്റെ compression ആണ്, magic അല്ല. ഒരു shallow network-നെക്കാൾ വളരെ കുറച്ച് total neuron-കൾ ഉപയോഗിച്ച് ഒരു deep network complex function-കൾ express ചെയ്യാം — പക്ഷേ ഓരോ extra layer-ഉം training signal അതിലൂടെ travel back ചെയ്യേണ്ട മറ്റൊരു layer ആണ്, ശരിയായ trick-കൾ വരുന്നത് വരെ deep network-കൾ train ചെയ്യാൻ ഇത്ര hard ആയത് കൃത്യമായി ഇതുകൊണ്ടാണ്.

Activation Function-കൾ: Nonlinear Secret Sauce

ജനങ്ങളെ surprise ചെയ്യുന്ന ഒരു fact ഇവിടെയുണ്ട്: activation function-കൾ ഇല്ലാതെ, depth worthless ആണ്. ഓരോ layer-ഉം ഒരു plain weighted sum ആയിരുന്നാൽ, പത്ത് layer-കൾ stack ചെയ്യുന്നത് ഒരു single layer-ന് mathematically identical ആയിരിക്കും — linear operation-കളുടെ ഒരു chain ഒരു linear operation-ലേക്ക് collapse ചെയ്യുന്നു. ഒരു glorified straight line build ചെയ്യാൻ compute-ൽ ഒരു fortune ചെലവാക്കിയിരിക്കും. Activation function-കൾ layer-കൾക്കിടയിൽ insert ചെയ്ത nonlinear kink ആണ്, network-നെ bend ചെയ്യാൻ അനുവദിക്കുന്നു, bending ആണ് അതിനെ powerful ആക്കുന്നത്.

ഓരോ practitioner-ഉം അറിഞ്ഞിരിക്കേണ്ട activation-കൾ:

  • Sigmoid — σ(x) = 1/(1+e⁻ˣ). ഏത് input-ഉം (0, 1)-ലേക്ക് squash ചെയ്യുന്നു, ഇത് ഒരു probability ആയി naturally read ചെയ്യുന്നു. ഒരു binary output-ന് great, പക്ഷേ hidden layer-കൾക്ക് അകത്ത് ഒരു poor choice: large positive അല്ലെങ്കിൽ negative input-കൾക്ക് curve flat ആകുന്നു, gradient almost zero-ലേക്ക് പോകുന്നു, learning halt ആകുന്നു — vanishing gradient problem.
  • Tanh — (−1, 1)-ൽ output ചെയ്യുന്നു, zero-centered ആണ്, optimization സഹായിക്കുന്നു. Extreme-കളിൽ ഇപ്പോഴും saturate ചെയ്യുന്നു, sigmoid-നെക്കാൾ കുറച്ച് harshly.
  • ReLU — f(x) = max(0, x). Hidden layer-കൾക്ക് modern default. Compute ചെയ്യാൻ laughably simple-ഉം cheap-ഉം ആണ്, എല്ലാ positive input-കൾക്കും ഒരു healthy gradient keep ചെയ്യുന്നു. അതിന്റെ flaw "dying ReLU" ആണ്: always zero output ചെയ്യാൻ pushed ചെയ്ത ഒരു neuron permanently inert ആയി മാറാം.
  • Leaky ReLU — f(x) = max(0.01x, x). negative-കൾക്ക് ഒരു small trickle of gradient through ചെയ്യാൻ അനുവദിക്കുന്നു, neuron-കൾ die ചെയ്യുന്നത് നിർത്തുന്നു.
  • GELU — BERT, GPT പോലുള്ള transformer-കൾക്ക് inside standard ആയ ReLU-ന്റെ ഒരു smooth, probabilistic cousin.
പ്രായോഗിക rule: ഏകദേശം എല്ലാ hidden layer-കൾക്കും ReLU (അല്ലെങ്കിൽ ഒരു variant) ഉപയോഗിക്കുക. ഒരു single yes/no output-ന് sigmoid ഉപയോഗിക്കുക, several class-കളിൽ ഒന്ന് പിക്ക് ചെയ്യാൻ softmax ഉപയോഗിക്കുക, predicted price പോലുള്ള regression output-കൾക്ക് activation ഇല്ല (linear) ഉപയോഗിക്കുക. Financial ML-ൽ നിങ്ങൾ build ചെയ്യുന്ന ഭൂരിഭാഗവും ഇത് കവർ ചെയ്യുന്നു.

Backpropagation-ഉം Gradient Descent-ഉം: Network-കൾ Learn ചെയ്യുന്നത് എങ്ങനെ

ഇതുവരെ, weight-കൾ നിറഞ്ഞ ഒരു network ഞങ്ങൾക്കുണ്ട്. പക്ഷേ ശരിയായ weight value-കൾ എവിടെ നിന്ന് വരുന്നു? ആരും type ചെയ്യുന്നില്ല അവയെ — ദശലക്ഷക്കണക്കിന് ഉണ്ട്. അവ ഒരു feedback loop-ലൂടെ discover ചെയ്യപ്പെടുന്നു, calculus bookkeeping ചെയ്യുന്ന organized trial and error ആണ് heart-ൽ.

Step one forward pass ആണ്: ഒരു example feed ചെയ്യുക, layer-കളിലൂടെ അത് flow ചെയ്യാൻ അനുവദിക്കുക, ഒരു prediction read off ചെയ്യുക. Step two ആ prediction എത്ര wrong ആയിരുന്നു എന്ന് ഒരു loss function ഉപയോഗിച്ച് measure ചെയ്യുകയാണ്. ഒരു number (ഒരു price) predict ചെയ്യാൻ, Mean Squared Error prediction-ഉം truth-ഉം തമ്മിലുള്ള squared gap average ചെയ്യുന്നു. Classification-ന്, cross-entropy predicted probabilities correct answer-ൽ നിന്ന് എത്ര ദൂരെയാണ് measure ചെയ്യുന്നു. High loss "very wrong" എന്നാണ് അർത്ഥം; training-ന്റെ entire goal ആ number താഴെ push ചെയ്യുകയാണ്.

Step three clever part ആണ്: backpropagation. Calculus-ന്റെ chain rule ഉപയോഗിച്ച്, algorithm loss-ൽ നിന്ന് ഓരോ layer-ലൂടെയും backward work ചെയ്യുന്നു, ഓരോ individual weight-ന്റെയും ഒരു gradient compute ചെയ്യുന്നു — "ഈ ഒരു weight ഞാൻ ഒരു ചെറിയ hair-ന് നുഡ്ജ് ചെയ്താൽ, error മുകളിലേക്ക് പോകുമോ താഴേക്ക് പോകുമോ, എത്ര?" എന്ന ചോദ്യത്തിന്റെ ഉത്തരം. AI winter thaw ചെയ്ത trick ഇതാണ്. 1980-കളിൽ popularize ചെയ്ത Backprop, guess ചെയ്യുന്നതിനുപകരം ദശലക്ഷക്കണക്കിന് weight-കൾക്കിടയിൽ efficiently blame assign ചെയ്യാൻ അതിനെ possible ആക്കി.

Step four gradient descent ആണ്: w = w − α × ∂L/∂w rule ഉപയോഗിച്ച്, loss കുറയ്ക്കുന്ന direction-ൽ ഓരോ weight-ഉം ഒരു small step nudge ചെയ്യുക. Valley floor-ലെത്താൻ ശ്രമിക്കുന്ന ഒരു foggy hillside-ൽ standing picture ചെയ്യുക; bottom നിങ്ങൾക്ക് കാണാൻ കഴിയില്ല, പക്ഷേ നിങ്ങളുടെ കാൽക്കീഴിൽ ഏത് വഴിയാണ് downhill എന്ന് feel ചെയ്യാൻ കഴിയും, അതിനാൽ ആ വഴി ഒരു step എടുത്ത് repeat ചെയ്യുക. Learning rate α നിങ്ങളുടെ stride length ആണ്, whole system-ലെ arguably most important dial ആണ് ഇത്. Too large ആയാൽ valley bound past ചെയ്ത് forever bounce ചെയ്യും; too small ആയാൽ eternity-ക്ക് inch along ചെയ്യും അല്ലെങ്കിൽ halfway down ഒരു ditch-ൽ stuck ആകും.

നിങ്ങൾ immediately meet ചെയ്യുന്ന രണ്ട് refinement-കൾ. Mini-batch-കൾ: ഓരോ single example-ന് ശേഷവും update ചെയ്യുന്നതിന് പകരം (noisy) അല്ലെങ്കിൽ entire dataset-ന് ശേഷം മാത്രം (slow), ഒരു small batch-ന് ശേഷം update ചെയ്യുന്നു — typically 32 to 256 example-കൾ — ഇത് stable-ഉം, GPU-friendly-ഉം, practice-ൽ sweet spot-ഉം ആണ്. Adam: ഓരോ weight-ന്റെയും adaptive stride നൽകി, noisy gradient-കൾ smooth ചെയ്യാൻ momentum ചേർക്കുന്ന ഒരു smarter optimizer. ഏകദേശം ഏത് project-ലും ആരംഭിക്കാൻ Adam sensible default ആണ് — എങ്കിലും ധനകാര്യത്തിൽ, low signal-to-noise ratio-ഓടെ, learning rate, weight decay, stopping point hand ഉപയോഗിച്ച് tune ചെയ്യേണ്ടതുണ്ട്.

Training loop: forward → loss → backward → update Forward activations Loss L MSE, CE… Backward ∂L/∂w w ← w−η∇ Mini-batches: noisy gradients, faster steps, better GPU use Adam: per-parameter learning rates + momentum (default starting point) Still tune η, weight decay, early stopping for finance (low SNR)
Autodiff chain rule walk ചെയ്യുന്നു; optimizer-കൾ എത്ര aggressive ആയി step ചെയ്യണം എന്ന് തീരുമാനിക്കുന്നു.

Overfitting: Network Learn ചെയ്യുന്നതിനുപകരം Memorize ചെയ്യുമ്പോൾ

ഇവിടെ ഏത് ധനകാര്യ model-നെക്കാളും കൂടുതൽ wreck ചെയ്യുന്ന failure mode ഉണ്ട്. ഒരു neural network ഇത്ര flexible ഒരു function-fitter ആണ്, chance ഉണ്ടെങ്കിൽ, അത് നിങ്ങളുടെ training ഡേറ്റ outright memorize ചെയ്യും — random noise, one-off flukes, ഒരിക്കലും repeat ചെയ്യാത്ത coincidence-കൾ ഉൾപ്പെടെ. അത് കണ്ട ഡേറ്റയിൽ beautifully score ചെയ്യുന്നു, പുതിയ എന്തെങ്കിലും meet ചെയ്യുന്ന നിമിഷം flat ആയി വീഴുന്നു. ഇതാണ് overfitting, market-കളിൽ — genuine signal faint ഉം noise deafening-ഉം ആയിടത്ത് — ഇത് default outcome ആണ്, exception അല്ല. ഇതിനെ fight ചെയ്യുന്ന tool-കളെ regularization എന്ന് വിളിക്കുന്നു.

Dropout ഏറ്റവും popular ആണ്. ഓരോ training step-ലും, ഓരോ neuron-ഉം temporarily switch off ചെയ്യപ്പെടാൻ ഒരു probability p ഉണ്ട് (പലപ്പോഴും 0.2–0.5). Random teammate-കൾ disappear ചെയ്ത് കൊണ്ടിരിക്കുമ്പോൾ പോലും work ചെയ്യാൻ force ചെയ്യപ്പെട്ട network, fragile ആയ കുറച്ചെണ്ണത്തിൽ എല്ലാം bet ചെയ്യുന്നതിനുപകരം അതിന്റെ knowledge നിരവധി neuron-കളിലുടനീളം spread ചെയ്യാൻ learn ചെയ്യുന്നു — ഒരു single model-ൽ baked ചെയ്ത ഒരു ensemble effect. Prediction time-ൽ എല്ലാ neuron-കളും appropriately scale ചെയ്ത് തിരികെ switch ചെയ്യുന്നു.

L1, L2 penalty-കൾ large weight-കൾക്ക് loss-ലേക്ക് ഒരു tax ചേർക്കുന്നു. L2 (weight decay) ഏത് single weight-ഉം too big ആകുന്നത് discourage ചെയ്യുന്നു, influence evenly spread ചെയ്യുന്നു. L1 sharper ആണ്: ഇത് useless weight-കൾ exactly zero-ലേക്ക് actively drive ചെയ്യുന്നു, automatic feature selection perform ചെയ്യുന്നു. നിങ്ങളുടെ hundred feature-കളിൽ ഭൂരിഭാഗവും noise ആയ ഒരു financial dataset-ൽ, L1 signal carry ചെയ്യുന്ന handful quietly കണ്ടെത്താം.

Batch normalization ഓരോ layer-ന്റെയും output-കൾ ഒരു stable mean, variance-ലേക്ക് rescale ചെയ്യുന്നു, ഇത് training steady ആക്കുന്നു, higher learning rate-കൾ ഉപയോഗിക്കാൻ അനുവദിക്കുന്നു. Early stopping എല്ലാ guard-കളിലും simplest-ഉം reliable-ഉം ആണ്: ഒരു held-out validation set-ലെ error watch ചെയ്യുക, training error താഴെ വീണുകൊണ്ടിരിക്കുമ്പോൾ up creep ചെയ്യാൻ ആരംഭിക്കുന്ന നിമിഷം, നിർത്തുക — ആ crossover ആണ് network learning patterns-ൽ നിന്ന് memorizing noise-ലേക്ക് switch ചെയ്യുന്ന exact moment.

ട്രേഡർമാർക്കുള്ള honest takeaway: ഒരു backtest-ൽ perfect ആയി കാണപ്പെടുന്ന ഒരു model ഏകദേശം certainly overfit ചെയ്തിരിക്കുന്നു. Real test മോഡൽ training-ൽ ഒരിക്കലും touch ചെയ്യാത്ത ഒരു time period-ലെ performance ആണ്. നിങ്ങളുടെ network out-of-sample-ൽ ഒരു simple baseline beat ചെയ്യാൻ കഴിയില്ലെങ്കിൽ, impressive backtest ഒരു mirage ആയിരുന്നു — ഒരു edge അല്ല.

Pattern-കൾക്ക് CNN-കൾ, Sequence-കൾക്ക് RNN-കളും LSTM-കളും

Plain feedforward network-കൾ ഓരോ input-ഉം unstructured numbers-ന്റെ ഒരു bag ആയി treat ചെയ്യുന്നു. പക്ഷേ ധാരാളം real ഡേറ്റക്ക് structure ഉണ്ട് — ഒരു image-ന് spatial layout ഉണ്ട്, ഒരു price series-ന് temporal order ഉണ്ട് — രണ്ട് specialized architecture-കൾ ആ structure directly exploit ചെയ്യുന്നു.

Convolutional Neural Network-കൾ (CNN-കൾ) image-കൾക്ക് build ചെയ്തതാണ്. ഓരോ pixel-ഉം ഓരോ neuron-ഉമായി connect ചെയ്യുന്നതിനുപകരം, ഒരു CNN input-നു ചുറ്റും ചെറിയ filter-കൾ slide ചെയ്യുന്നു, local pattern-കൾക്കായി — ഇവിടെ ഒരു edge, അവിടെ ഒരു texture — അതേ filter എല്ലായിടത്തും reuse ചെയ്ത് hunt ചെയ്യുന്നു. Early layer-കൾ edge-കൾ catch ചെയ്യുന്നു; deeper layer-കൾ shapes, object-കളാക്കി അവയെ assemble ചെയ്യുന്നു. ധനകാര്യത്തിൽ, researcher-മാർ CNN-കൾ ഇവയിലേക്ക് point ചെയ്തിട്ടുണ്ട്:

  • Candlestick chart image-കൾ — chart reading-നെ ഒരു visual pattern-recognition ജോലിയായി reframe ചെയ്യുന്നു.
  • Order-book heatmap-കൾ — depth-of-market snapshot-കളിലെ supply/demand imbalance-കൾ spot ചെയ്യുന്നു.
  • Raw price sequence-കൾക്ക് 1D convolution-കൾ — ഒരു image CNN spatial motif-കൾ learn ചെയ്യുന്ന രീതിയിൽ short temporal motif-കൾ learn ചെയ്യുന്നു.

Recurrent Neural Network-കൾ (RNN-കൾ) sequence-കൾക്ക് build ചെയ്തതാണ്. അവ ഒരു hidden state — ഒരു running memory — ഓരോ step-ൽ നിന്നും next-ലേക്ക് carry ചെയ്ത്, ഒരു സമയത്ത് ഒരു time step process ചെയ്യുന്നു. Theory-യിൽ ഇത് time series-ന് ideal ആക്കുന്നു. Practice-ൽ, vanilla RNN-കൾ long sequence-കളിലുടനീളം vanishing gradient problem കൊണ്ട് crippled ആണ്: many step-കൾക്ക് മുമ്പുള്ള signal, back വരുന്ന വഴിയിൽ zero-ലേക്ക് shrink ചെയ്യുന്നു, അതിനാൽ network distant past simply forget ചെയ്യുന്നു.

Long Short-Term Memory (LSTM) network-കൾ, ഓരോ step-ലും memory-യിൽ നിന്ന് എന്ത് erase ചെയ്യണം, എന്ത് write in ചെയ്യണം, എന്ത് read out ചെയ്യണം എന്ന് decide ചെയ്യുന്ന ഒരു small set of learnable gate-കൾ — forget, input, output — ഉപയോഗിച്ച് ഇത് fix ചെയ്യുന്നു. ആ gating-ന് ഒരു LSTM-നെ ഹൻഡ്രഡ് step-കൾക്ക് ഉടനീളം relevant information hold ചെയ്യാൻ അനുവദിക്കുന്നു, ഇന്നലത്തെ ഒരു event ഇന്നത്തെ price ഇപ്പോഴും move ചെയ്യുന്ന financial time series model ചെയ്യാൻ ഇത് ഒരു workhorse ആയത് ഇതുകൊണ്ടാണ്.

import torch.nn as nn

class PricePredictor(nn.Module):
    def __init__(self, input_dim, hidden_dim, num_layers):
        super().__init__()
        self.lstm = nn.LSTM(input_dim, hidden_dim,
                            num_layers, batch_first=True,
                            dropout=0.2)
        self.fc = nn.Linear(hidden_dim, 1)

    def forward(self, x):
        out, _ = self.lstm(x)
        return self.fc(out[:, -1, :])

WebSocket feed-കൾ continuous price, order-book ഡേറ്റ stream ചെയ്യുന്ന GaiaEx പോലുള്ള ഒരു platform-ൽ, ഒരു LSTM ആ sequence naturally ingest ചെയ്യാം. അതു പറഞ്ഞെങ്കിലും, field ഭൂരിഭാഗം moved on ചെയ്തിരിക്കുന്നു: modern language model-കൾക്ക് പിന്നിലുള്ള architecture ആയ transformer-കൾ — ഓരോ time step-ലൂടെയും ഒന്നൊന്നായി marching ചെയ്യുന്നതിനുപകരം ഒരേസമയം എല്ലാ time step-കളും look ചെയ്യാൻ "attention" ഉപയോഗിച്ച് നിരവധി sequence task-കളിൽ LSTM-കളെ ഇപ്പോൾ match ചെയ്യുകയോ beat ചെയ്യുകയോ ചെയ്യുന്നു.

Neural Network-കൾ Fix ചെയ്യാത്തത് (പ്രത്യേകിച്ച് Market-കളിൽ)

Neural network-കൾ spectacular pattern-finder-കൾ ആണ്. പക്ഷേ honest education എന്നാൽ അവ break ചെയ്യുന്നിടം പേരിടുന്നതാണ് — ട്രേഡിംഗിൽ, അവ expensive രീതികളിൽ break ചെയ്യുന്നു.

  • ഒരു pattern-ഉം exist ചെയ്യാത്തിടത്ത് പോലും അവ pattern-കൾ കണ്ടെത്തുന്നു. ഒരു deep network-ന് noise നൽകുക, അത് confidently fit ചെയ്യും. Market-കൾ ഭൂരിഭാഗവും noise ആണ്, അതിനാൽ ഒരു pattern real-ഉം persistent-ഉം ആണ് എന്ന് കാണിക്കാനുള്ള proof-ന്റെ burden നിങ്ങളുടേതാണ് — last year's ഡേറ്റയിൽ ഉള്ളത് മാത്രമല്ല.
  • അവ black box-കൾ ആണ്. ഒരു മോഡൽ accurate ആയിരിക്കുകയും എന്തുകൊണ്ട് എന്ന് നിങ്ങളോട് പറയാൻ unable ആയിരിക്കുകയും ചെയ്യാം. അത് suddenly പണം നഷ്ടപ്പെടുത്തുമ്പോൾ, പലപ്പോഴും clean explanation-ഉം obvious fix-ഉം ഇല്ല. Risk-bearing capital-ന്, "അത് അങ്ങനെ ചെയ്ത കാരണം എനിക്കറിയില്ല" ഒരു serious liability ആണ്.
  • അവ hungry-ഉം fragile-ഉം ആണ്. ദശലക്ഷക്കണക്കിന് clean example-കൾ ഉള്ളപ്പോൾ deep learning shine ചെയ്യുന്നു. Financial history short, non-stationary, noisy ആണ് — നിങ്ങളുടെ model train ചെയ്ത regime simply exist ചെയ്യുന്നത് നിർത്താം, distribution shift എന്ന ഒരു problem.
  • അവ free ആയി simpler model-കളെ beat ചെയ്യുന്നില്ല. ധനകാര്യത്തിൽ സാധാരണമായ tabular, feature-based ഡേറ്റയിൽ, XGBoost പോലുള്ള gradient-boosted tree-കൾ പലപ്പോഴും neural network-കളെ outperform ചെയ്യുന്നു, train ചെയ്യാൻ faster-ഉം interpret ചെയ്യാൻ easier-ഉം ആയിരിക്കുമ്പോൾ. Complexity ഒരു cost ആണ്, ഒരു virtue അല്ല.
  • അവ attack ചെയ്യപ്പെടാം. ഒരു input-ലേക്കുള്ള ചെറിയ, deliberate perturbation-കൾ ഒരു network-ന്റെ output flip ചെയ്യാം — ഒരു adversarial example — നിങ്ങളുടെ മോഡൽ കാണുന്ന കാര്യങ്ങൾ ഒരു adversary-ക്ക് shape ചെയ്യാൻ കഴിയുന്ന എവിടെയും ഇത് പ്രധാനമാണ്.
Winner-കളെയും blowup-കളെയും വേർതിരിക്കുന്ന discipline: ഒരു neural network-നെ എപ്പോഴും ഒരു simple baseline-നെതിരെ benchmark ചെയ്യുക, training-ൽ മോഡൽ ഒരിക്കലും കണ്ടിട്ടില്ലാത്ത ഒരു period-ലെ ഡേറ്റയിൽ എപ്പോഴും evaluate ചെയ്യുക, stress-test ചെയ്യാൻ കഴിയാത്ത ഒരു മോഡലിന് പിന്നിൽ ഒരിക്കലും capital deploy ചെയ്യരുത്. AlphaGo-ക്ക് ഒരു perfect simulator-ഉം unlimited self-play-ഉം ഉണ്ടായിരുന്നു. Market ഒരു noisy, non-repeating history നൽകുന്നു — വ്യത്യാസം respect ചെയ്യുക.

Hardware, Tool-കൾ, നിങ്ങളുടെ മുന്നോട്ടുള്ള Path

Neural network-കൾ train ചെയ്യുന്നത് arithmetic-ൽ heavy ആണ് — ഭൂരിഭാഗവും matrix multiplication-കൾ — നിങ്ങൾ അത് ഏത് hardware-ൽ run ചെയ്യുന്നു എന്നത് നിങ്ങളുടെ iteration speed "overnight"-ൽ നിന്ന് "over coffee"-ലേക്ക് മാറ്റുന്നു.

GPU-കൾ standard ആണ്. അവയുടെ ആയിരക്കണക്കിന് core-കൾ ഒരേസമയം വ്യത്യസ്ത ഡേറ്റയിൽ അതേ operation run ചെയ്യുന്നു, ഇത് neural network math-ന്റെ exact shape ആണ്. CPU-യിൽ 8 മണിക്കൂർ crawl ചെയ്യുന്ന ഒരു മോഡൽ ഒരു modern GPU-യിൽ 15 മിനിറ്റിനുള്ളിൽ finish ചെയ്യാം. NVIDIA dominate ചെയ്യുന്നു: consumer card-കൾ (ഒരു RTX-class GPU) learning-ന് fine ആണ്, data-center card-കൾ (A100, H100), cloud instance-കൾ production-scale training handle ചെയ്യുന്നു. Google-ന്റെ tensor math-ന്റെ custom chip-കൾ ആയ TPU-കൾ, Google Cloud വഴി very large model-കളിൽ shine ചെയ്യുന്നു, പക്ഷേ broader software support-ന് GPU-കൾ മിക്ക practitioner-മാർക്കും practical default ആയി തുടരുന്നു.

ആരംഭിക്കാൻ നിങ്ങൾ ഒന്നും buy ചെയ്യേണ്ടതില്ല. Google Colab learning-നും prototyping-നും മതിയായ free GPU time hand out ചെയ്യുന്നു; നിങ്ങളുടെ model-കൾ growth ചെയ്യുന്നതനുസരിച്ച് on demand cloud GPU-കൾ rent ചെയ്യുക; നിങ്ങൾ daily train ചെയ്യുന്നപ്പോളും cloud bill justify ചെയ്യുന്നപ്പോളും മാത്രം dedicated hardware buy ചെയ്യുക.

ഇവിടെ നിന്ന് ഒരു sane learning path:

  • Few engineered feature-കളിൽ ഒരു plain feedforward network-ൽ ആരംഭിക്കുക — fast to train, easy to debug, ഒരു honest baseline.
  • Raw price sequence-കളിൽ ഒരു 1D CNN try ചെയ്യുക, ആ baseline-നോട് head-to-head compare ചെയ്യുക.
  • GaiaEx-ന്റെ API-യിൽ നിന്ന് candlestick ഡേറ്റയുടെ sequence-കൾ consume ചെയ്യുന്ന ഒരു LSTM build ചെയ്യുക.
  • Attention-ഉം transformer-കളും study ചെയ്യുക — sequence-കൾക്ക് increasingly architecture of choice ആണ്.
  • ഏറ്റവും പ്രധാനമായി, നിങ്ങളുടെ deep മോഡലിനെ എപ്പോഴും XGBoost പോലുള്ള ഒരു gradient-boosting baseline-നെതിരെ pit ചെയ്യുക. Neural network-ന് നിങ്ങളുടെ tabular financial ഡേറ്റയിൽ out-of-sample-ൽ simpler മോഡലിനെ beat ചെയ്യാൻ കഴിഞ്ഞില്ലെങ്കിൽ, simpler മോഡൽ ship ചെയ്ത് move on ചെയ്യുക.

Move 37 കണ്ടെത്തിയ അതേ loop — predict ചെയ്യുക, error measure ചെയ്യുക, weight-കൾ nudge ചെയ്യുക, repeat ചെയ്യുക — ഈ ലിസ്റ്റിലെ ഓരോ system-നും ഉള്ളിൽ run ചെയ്യുന്ന loop ഇതാണ്. ആ loop മനസ്സിലാക്കുക, deep learning നിങ്ങൾ മനസ്സിലാക്കും; ബാക്കി architecture-ഉം engineering-ഉം ആണ് അതിന് മുകളിൽ.